Views
No views yet
Q4_0_ROCMFP4_STRIX_LEAN, ~4.29 bpw) GGUF quant of
Qwen/Qwen-AgentWorld-35B-A3B
— a 35B-A3B Mixture-of-Experts (qwen35moe) world-model, 262K context.llama.cpp or standard GGUF tooling —
those will fail with 101 is not a valid GGMLQuantizationType.ghcr.io/hal0ai/amd-strix-halo-toolboxes:rocm-7.2.4-rocmfp4-server, orrocmfp4-llama
fork (branch mtp-rocmfp4-strix), targeting gfx1151.| Architecture | qwen35moe (35B total / ~3B active, MoE) |
| Quant | Q4_0_ROCMFP4_STRIX_LEAN (~4.38 bpw target; 4.29 bpw measured) |
| Recipe | ROCmFP4 experts/FFN + Strix K/V + Q5_K token embeddings |
| File size | ~17.3 GiB (from 66 GiB BF16) |
| MTP | No (plain quant, no speculative-decode draft head) |
| Context | 262144 |
| Base | Qwen/Qwen-AgentWorld-35B-A3B (Apache-2.0) |
imatrix_unsloth.gguf), via the hal0
ROCmFP4 quantize pipeline (llama-quantize --imatrix ... Q4_0_ROCMFP4_STRIX_LEAN).