Views
No views yet
Quantized release
This repository is an AutoRound W4A16 quantized version ofdeepreinforce-ai/Ornith-1.0-397B.
It is not the original BF16 model repository.
| Item | Value |
|---|---|
| Base model | deepreinforce-ai/Ornith-1.0-397B |
| Quantized repo | Pilcothink/Ornith-1.0-397B-W4A16-AutoRound |
| Quantization method | AutoRound |
| Quantization format | W4A16 |
| Weight precision | 4-bit |
| Activation precision | 16-bit |
| Base model relation | quantized |
| License | MIT, inherited from upstream model |
The following section is copied from the upstream model card ofdeepreinforce-ai/Ornith-1.0-397B.
It describes the original BF16 model, not this AutoRound W4A16 quantized release.


| Ornith-1.0-397B | Qwen3.5-397B | Qwen3.7-Max | GLM-5.2-744B | Minimax-M3-428B | DeepSeek-V4-Pro-1.6T | Claude Opus 4.7 | Claude Opus 4.8 | |
|---|---|---|---|---|---|---|---|---|
| Agentic Coding | ||||||||
| Terminal-Bench 2.1 (Terminus-2) | 77.5 | 53.5 | 73.5 | 81.0 | 64 | 64 | 70.3 | 85 |
| Terminal-Bench 2.1 (Claude Code) | 78.2 | 48.6 | 69.8 | 82.7 | - | 66.5 | 69.7 | 78.9 |
| SWE-bench Verified | 82.4 | 76.4 | 80.4 | - | - | 80.6 | 80.8 | 87.6 |
| SWE-bench Pro | 62.2 | 51.6 | 60.6 | 62.1 | 59 | 55.4 | 64.3 | 69.2 |
| SWE-bench Multilingual | 78.9 | 69.3 | 78.3 | - | - | 76.2 | - | - |
| NL2Repo | 48.2 | 36.8 | 47.2 | 48.9 | 42.1 | - | - | 69.7 |
| Claw-eval Avg | 77.1 | 70.7 | 65.2 | - | - | 75.8 | 78.2 | - |
| SWE Atlas - QnA | 41.2 | 20.4 | - | - | 37.9 | 27.2 | 40.3 | 48.8 |
| SWE Atlas - RF | 42.6 | 18.4 | - | - | - | - | 48.6 | 46.7 |
| SWE Atlas - TW | 39.1 | 18.5 | - | - | 30.8 | - | 38.5 | - |
<think> … </think> block before the final answer. The serving recipes below enable a reasoning parser so the chain-of-thought is returned in a separate reasoning_content field, and a tool-call parser so the model's <tool_call> blocks are surfaced as OpenAI-style tool_calls.--tensor-parallel-size / --tp to the number of GPUs you have.1vllm serve deepreinforce-ai/Ornith-1.0-397B \
2 --served-model-name Ornith-1.0-397B \
3 --tensor-parallel-size 8 \
4 --host 0.0.0.0 --port 8000 \
5 --max-model-len 262144 \
6 --gpu-memory-utilization 0.90 \
7 --enable-prefix-caching \
8 --enable-auto-tool-choice --tool-call-parser qwen3_xml \
9 --reasoning-parser qwen3 \
10 --trust-remote-code1python -m sglang.launch_server \
2 --model-path deepreinforce-ai/Ornith-1.0-397B \
3 --served-model-name Ornith-1.0-397B \
4 --tp 8 \
5 --host 0.0.0.0 --port 8000 \
6 --context-length 262144 \
7 --mem-fraction-static 0.85 \
8 --tool-call-parser qwen3_coder \
9 --reasoning-parser qwen3transformers >= 5.8.1.1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_name = "deepreinforce-ai/Ornith-1.0-397B"
4
5tokenizer = AutoTokenizer.from_pretrained(model_name)
6model = AutoModelForCausalLM.from_pretrained(
7 model_name,
8 dtype="auto",
9 device_map="auto",
10)
11
12messages = [
13 {"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}
14]
15text = tokenizer.apply_chat_template(
16 messages,
17 tokenize=False,
18 add_generation_prompt=True,
19)
20
21inputs = tokenizer(text, return_tensors="pt").to(model.device)
22generated = model.generate(
23 **inputs,
24 max_new_tokens=512,
25 do_sample=True,
26 temperature=0.6,
27 top_p=0.95,
28 top_k=20,
29)
30output_ids = generated[0][inputs.input_ids.shape[1]:]
31
32# The reply contains a <think> ... </think> reasoning block followed by the answer.
33content = tokenizer.decode(output_ids, skip_special_tokens=True)
34print(content)</think> marker:1text = tokenizer.decode(output_ids, skip_special_tokens=True)
2if "</think>" in text:
3 reasoning, answer = text.split("</think>", 1)
4 reasoning = reasoning.replace("<think>", "").strip()
5 answer = answer.strip()
6else:
7 reasoning, answer = "", text.strip()1from openai import OpenAI
2
3client = OpenAI(
4 base_url="http://localhost:8000/v1",
5 api_key="EMPTY", # any non-empty string works for a local server
6)
7
8response = client.chat.completions.create(
9 model="Ornith-1.0-397B",
10 messages=[
11 {"role": "user", "content": "Write a one-line Python lambda that squares a number."}
12 ],
13 temperature=0.6,
14 top_p=0.95,
15 max_tokens=1024,
16)
17
18message = response.choices[0].message
19# reasoning_content holds the <think> trace; content holds the final answer.
20print("reasoning:", getattr(message, "reasoning_content", None))
21print("answer:", message.content)tool_calls field:1tools = [
2 {
3 "type": "function",
4 "function": {
5 "name": "get_weather",
6 "description": "Get the current weather for a city",
7 "parameters": {
8 "type": "object",
9 "properties": {"city": {"type": "string"}},
10 "required": ["city"],
11 },
12 },
13 }
14]
15
16response = client.chat.completions.create(
17 model="Ornith-1.0-397B",
18 messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
19 tools=tools,
20 tool_choice="auto",
21 temperature=0.6,
22 max_tokens=2048,
23)
24
25tool_call = response.choices[0].message.tool_calls[0]
26print(tool_call.function.name, tool_call.function.arguments)
27# -> get_weather {"city": "Paris"}curl at the same /v1/chat/completions endpoint.1import os
2from openai import OpenAI
3
4client = OpenAI(
5 base_url=os.getenv("OPENAI_BASE_URL", "http://localhost:8000/v1"),
6 api_key=os.getenv("OPENAI_API_KEY", "EMPTY"),
7)
8
9tools = [
10 {
11 "type": "function",
12 "function": {
13 "name": "run_shell",
14 "description": "Run a shell command and return its output.",
15 "parameters": {
16 "type": "object",
17 "properties": {
18 "command": {"type": "string", "description": "The command to run"}
19 },
20 "required": ["command"],
21 },
22 },
23 }
24]
25
26messages = [{"role": "user", "content": "List the Python files in the current directory."}]
27
28response = client.chat.completions.create(
29 model="deepreinforce-ai/Ornith-1.0-397B",
30 messages=messages,
31 tools=tools,
32 temperature=0.6,
33 top_p=0.95,
34)
35print(response.choices[0].message)1# Hermes talks to any OpenAI-compatible endpoint — point it at your Ornith server.
2export OPENAI_BASE_URL="http://localhost:8000/v1"
3export OPENAI_API_KEY="EMPTY"
4export MODEL="deepreinforce-ai/Ornith-1.0-397B"1# OpenClaw talks to any OpenAI-compatible endpoint — point it at your Ornith server.
2export OPENAI_BASE_URL="http://localhost:8000/v1"
3export OPENAI_API_KEY="EMPTY"
4export OPENAI_MODEL="deepreinforce-ai/Ornith-1.0-397B"1pip install unsloth
2
3# Load Ornith for fast local inference or fine-tuning (Python):
4# from unsloth import FastLanguageModel
5# model, tokenizer = FastLanguageModel.from_pretrained(
6# "deepreinforce-ai/Ornith-1.0-397B",
7# max_seq_length=262144,
8# load_in_4bit=True,
9# )1pip install openhands-ai
2
3# OpenHands routes through LiteLLM; the "openai/" prefix selects the OpenAI-compatible path.
4export LLM_MODEL="openai/deepreinforce-ai/Ornith-1.0-397B"
5export LLM_BASE_URL="http://localhost:8000/v1"
6export LLM_API_KEY="EMPTY"
7
8# Launch the CLI (or run the official OpenHands Docker image with the same env vars).
9openhandsOPENAI_BASE_URL and OPENAI_API_KEY) to understand large codebases, automate tedious work, and ship faster.1# Register your local Ornith endpoint as a provider in ~/.config/opencode/opencode.json:
2#
3# {
4# "$schema": "https://opencode.ai/config.json",
5# "provider": {
6# "ornith": {
7# "npm": "@ai-sdk/openai-compatible",
8# "name": "Ornith (local)",
9# "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "EMPTY" },
10# "models": { "deepreinforce-ai/Ornith-1.0-397B": { "name": "Ornith-1.0-397B" } }
11# }
12# }
13# }
14
15opencode1@misc{ornith_397b,
2 title = {{Ornith-1.0-397B}: Agentic Coding, Open to All},
3 url = {https://deep-reinforce.com/ornith_1_0.html},
4 author = {{DeepReinforce Team}},
5 year = {2026}
6}