Views
No views yet
| File | Base | Context | Notes |
|---|---|---|---|
gguf/qwen35-claude-coder-4b.gguf | Qwen3.5 4B | 64K | Light, fast agent for 16GB Apple Silicon. ~30 tok/s. |
gguf/qwen35-claude-coder-9b.gguf | Qwen3.5 9B | 64K | Stronger, production-quality code. ~17 tok/s on 32GB, ~14 on 16GB. |
ollama run rafw007/qwen35-claude-coder:9b
ollama launch claude --model rafw007/qwen35-claude-coder:9b*-mlx) exist ONLY inside a local Ollama install and were tested ONLY there.
They are stored in Ollama internal MLX format (nvfp4) and were not pushed to the ollama.com registry, which currently rejects MLX-format manifests. They are not provided here as standalone mlx_lm weights and were not validated outside Ollama. This HF repo ships the portable GGUF weights plus the Modelfiles (full recipe) for every variant, including the MLX ones, so the build is reproducible. The published, downloadable models are the GGUF ones on ollama.com (rafw007/qwen35-claude-coder:4b and :9b).modelfiles/. Sampling: temperature 0.2, top_p 0.9, top_k 20, repeat_penalty 1.05, num_ctx 65536. System prompt enforces: act with tools now, write files, ground in real output, be terse, one language, never drift to Chinese.