Views
No views yet
Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled model, optimized natively for Apple Silicon using the Asgard AI Platform (Heimdall) infrastructure.| Task Type | Context (Max Tokens) | TTFT (Time to First Token) | TPS (Tokens per Second) |
|---|---|---|---|
| Short (Chat) | 20 | 1.62s - 3.51s | ~12.5 tok/s |
| Medium (RAG) | 200 | ~13.20s | ~15.1 tok/s |
| Long (Reasoning) | 500 | ~32.65s | ~15.3 tok/s |
~/Developer/Heimdall/reports/benchmark_20260402_140737.html1cd ~/Developer/Heimdall
2LLM_MODEL="$HOME/Developer/Heimdall/models/Qwen3.5-27B-Opus-Reasoning-MLX-4bit" ./scripts/start.sh