Text-only MXFP4 MLX build of
deepreinforce-ai/Ornith-1.0-35B, packaged for MTPLX native-MTP inference on Apple Silicon.
This is intended for local, private inference. After download, prompts and outputs can stay on your machine when served with a local MTPLX endpoint.
1python -m mtplx.server.openai \
2 --model /path/to/ornith-1.0-35b-mxfp4-mtplx \
3 --backend-id qwen3_next \
4 --generation-mode mtp \
5 --load-mtp \
6 --depth 2 \
7 --profile sustained \
8 --chat-template-profile tokenizer \
9 --normalize-thinking-tags \
10 --reasoning-mode on \
11 --enable-thinking \
12 --reasoning-parser qwen3 \
13 --reasoning-effort high \
14 --temperature 0.2 \
15 --top-p 0.95 \
16 --top-k 20 \
17 --no-stats-footer
Hardware reference: Apple M4 Max Apple Silicon with 64 GB unified memory.
On a local Apple Silicon host, this MTPLX profile matched the LM Studio text baseline on a small hard validation suite:
This is a lightweight local validation, not a public leaderboard result.
This repository contains model files only. It does not include a hosted endpoint, telemetry, or an external service requirement. Use a local server and inspect your client configuration if strict data locality matters.