Views
No views yet

| Ornith-1.0-9B | Qwen3.5-9B | Qwen3.5-35B | Gemma4-12B | Gemma4-31B | |
|---|---|---|---|---|---|
| Agentic Coding | |||||
| Terminal-Bench 2.1 (Terminus-2) | 43.1 | 21.3 | 41.4 | 21 | 42.1 |
| Terminal-Bench 2.1 (Claude Code) | 40.6 | 18.9 | 38.9 | - | - |
| SWE-bench Verified | 69.4 | 53.2 | 70 | 44.2 | 52 |
| SWE-bench Pro | 42.9 | 31.3 | 44.6 | 27.6 | 35.7 |
| SWE-bench Multilingual | 52 | 39.7 | 60.3 | 32.5 | 51.7 |
| NL2Repo | 27.2 | 16.2 | 20.5 | 10.3 | 15.5 |
| Claw-eval Avg | 63.1 | 53.2 | 65.4 | 32.5 | 48.5 |
| SWE Atlas - QnA | 17.9 | 9.2 | 13.2 | - | - |
| SWE Atlas - RF | 16.6 | 4.3 | 10.2 | - | - |
| SWE Atlas - TW | 15.3 | 4.4 | 9.8 | - | - |
<think> … </think> block before the final answer. The serving recipes below enable a reasoning parser so the chain-of-thought is returned in a separate reasoning_content field, and a tool-call parser so the model's <tool_call> blocks are surfaced as OpenAI-style tool_calls.temperature=0.6, top_p=0.95, top_k=20 (use temperature=1.0 to reproduce the reported benchmark setup).--tensor-parallel-size / --tp if you want to shard across more GPUs.1vllm serve deepreinforce-ai/Ornith-1.0-9B \
2 --served-model-name Ornith-1.0-9B \
3 --host 0.0.0.0 --port 8000 \
4 --max-model-len 262144 \
5 --gpu-memory-utilization 0.90 \
6 --enable-prefix-caching \
7 --enable-auto-tool-choice --tool-call-parser qwen3_xml \
8 --reasoning-parser qwen3 \
9 --trust-remote-code1python -m sglang.launch_server \
2 --model-path deepreinforce-ai/Ornith-1.0-9B \
3 --served-model-name Ornith-1.0-9B \
4 --host 0.0.0.0 --port 8000 \
5 --context-length 262144 \
6 --mem-fraction-static 0.85 \
7 --tool-call-parser qwen3_coder \
8 --reasoning-parser qwen3transformers >= 5.8.1.1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_name = "deepreinforce-ai/Ornith-1.0-9B"
4
5tokenizer = AutoTokenizer.from_pretrained(model_name)
6model = AutoModelForCausalLM.from_pretrained(
7 model_name,
8 dtype="auto",
9 device_map="auto",
10)
11
12messages = [
13 {"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}
14]
15text = tokenizer.apply_chat_template(
16 messages,
17 tokenize=False,
18 add_generation_prompt=True,
19)
20
21inputs = tokenizer(text, return_tensors="pt").to(model.device)
22generated = model.generate(
23 **inputs,
24 max_new_tokens=512,
25 do_sample=True,
26 temperature=0.6,
27 top_p=0.95,
28 top_k=20,
29)
30output_ids = generated[0][inputs.input_ids.shape[1]:]
31
32# The reply contains a <think> ... </think> reasoning block followed by the answer.
33content = tokenizer.decode(output_ids, skip_special_tokens=True)
34print(content)</think> marker:1text = tokenizer.decode(output_ids, skip_special_tokens=True)
2if "</think>" in text:
3 reasoning, answer = text.split("</think>", 1)
4 reasoning = reasoning.replace("<think>", "").strip()
5 answer = answer.strip()
6else:
7 reasoning, answer = "", text.strip()1from openai import OpenAI
2
3client = OpenAI(
4 base_url="http://localhost:8000/v1",
5 api_key="EMPTY", # any non-empty string works for a local server
6)
7
8response = client.chat.completions.create(
9 model="Ornith-1.0-9B",
10 messages=[
11 {"role": "user", "content": "Write a one-line Python lambda that squares a number."}
12 ],
13 temperature=0.6,
14 top_p=0.95,
15 max_tokens=1024,
16)
17
18message = response.choices[0].message
19# reasoning_content holds the <think> trace; content holds the final answer.
20print("reasoning:", getattr(message, "reasoning_content", None))
21print("answer:", message.content)tool_calls field:1tools = [
2 {
3 "type": "function",
4 "function": {
5 "name": "get_weather",
6 "description": "Get the current weather for a city",
7 "parameters": {
8 "type": "object",
9 "properties": {"city": {"type": "string"}},
10 "required": ["city"],
11 },
12 },
13 }
14]
15
16response = client.chat.completions.create(
17 model="Ornith-1.0-9B",
18 messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
19 tools=tools,
20 tool_choice="auto",
21 temperature=0.6,
22 max_tokens=2048,
23)
24
25tool_call = response.choices[0].message.tool_calls[0]
26print(tool_call.function.name, tool_call.function.arguments)
27# -> get_weather {"city": "Paris"}curl at the same /v1/chat/completions endpoint.1import os
2from openai import OpenAI
3
4client = OpenAI(
5 base_url=os.getenv("OPENAI_BASE_URL", "http://localhost:8000/v1"),
6 api_key=os.getenv("OPENAI_API_KEY", "EMPTY"),
7)
8
9tools = [
10 {
11 "type": "function",
12 "function": {
13 "name": "run_shell",
14 "description": "Run a shell command and return its output.",
15 "parameters": {
16 "type": "object",
17 "properties": {
18 "command": {"type": "string", "description": "The command to run"}
19 },
20 "required": ["command"],
21 },
22 },
23 }
24]
25
26messages = [{"role": "user", "content": "List the Python files in the current directory."}]
27
28response = client.chat.completions.create(
29 model="deepreinforce-ai/Ornith-1.0-9B",
30 messages=messages,
31 tools=tools,
32 temperature=0.6,
33 top_p=0.95,
34)
35print(response.choices[0].message)1# Hermes talks to any OpenAI-compatible endpoint — point it at your Ornith server.
2export OPENAI_BASE_URL="http://localhost:8000/v1"
3export OPENAI_API_KEY="EMPTY"
4export MODEL="deepreinforce-ai/Ornith-1.0-9B"1# Both runtimes load a GGUF build of Ornith (publish one at deepreinforce-ai/Ornith-1.0-9B-GGUF).
2
3# llama.cpp — serve an OpenAI-compatible API on port 8000.
4llama-server -hf deepreinforce-ai/Ornith-1.0-9B-GGUF --port 8000 -c 262144
5
6# Ollama — pull and chat with the same GGUF straight from Hugging Face.
7ollama run hf.co/deepreinforce-ai/Ornith-1.0-9B-GGUF1# OpenClaw talks to any OpenAI-compatible endpoint — point it at your Ornith server.
2export OPENAI_BASE_URL="http://localhost:8000/v1"
3export OPENAI_API_KEY="EMPTY"
4export OPENAI_MODEL="deepreinforce-ai/Ornith-1.0-9B"1pip install unsloth
2
3# Load Ornith for fast local inference or fine-tuning (Python):
4# from unsloth import FastLanguageModel
5# model, tokenizer = FastLanguageModel.from_pretrained(
6# "deepreinforce-ai/Ornith-1.0-9B",
7# max_seq_length=262144,
8# load_in_4bit=True,
9# )1pip install openhands-ai
2
3# OpenHands routes through LiteLLM; the "openai/" prefix selects the OpenAI-compatible path.
4export LLM_MODEL="openai/deepreinforce-ai/Ornith-1.0-9B"
5export LLM_BASE_URL="http://localhost:8000/v1"
6export LLM_API_KEY="EMPTY"
7
8# Launch the CLI (or run the official OpenHands Docker image with the same env vars).
9openhandsOPENAI_BASE_URL and OPENAI_API_KEY) to understand large codebases, automate tedious work, and ship faster.1# Register your local Ornith endpoint as a provider in ~/.config/opencode/opencode.json:
2#
3# {
4# "$schema": "https://opencode.ai/config.json",
5# "provider": {
6# "ornith": {
7# "npm": "@ai-sdk/openai-compatible",
8# "name": "Ornith (local)",
9# "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "EMPTY" },
10# "models": { "deepreinforce-ai/Ornith-1.0-9B": { "name": "Ornith-1.0-9B" } }
11# }
12# }
13# }
14
15opencode1@misc{ornith_9b,
2 title = {{Ornith-1.0-9B}: Agentic Coding, Open to All},
3 url = {https://deep-reinforce.com/ornith_1_0.html},
4 author = {{DeepReinforce Team}},
5 year = {2026}
6}