ChadrockV3 is the ROCmFP4 UE4M4 STRIX FAST build of Ornith-1.0-35B: a 34.66B-parameter, 256-expert / top-8-active Qwen3.5-MoE agentic coding model.
The goal is simple: keep legacy ROCmFP4-class throughput while spending ROCmFP4's previously unused scale high bit on a more accurate unsigned E4M4 scale. On the sampled held-out Ornith reconstruction gate, this reduced activation-weighted reconstruction error by 14.024% versus legacy UE4M3 ROCmFP4 FAST at the same 4.25 BPW block geometry. In full-model gfx1151 throughput measurements, ChadrockV3 delivered +2.04% PP512, +3.97% PP2048, and +0.21% TG128 mean throughput versus the near-size-matched legacy ROCmFP4 STRIX_LEAN control.
Required runtime: this GGUF uses private UE4M4 type 108 / ftype 121 and will not load in stock llama.cpp or legacy ROCmFPX. Build the public hstolte11-collab/ROCmFPX fork from its default ue4m4-fast branch. Complete build and launch commands are below.
Important: the 14.024% result is a sampled, activation-weighted tensor reconstruction result—not complete-model perplexity or a downstream task score. Comparative full-model PPL was not completed, so this release does not claim a 14% model-quality improvement.
Why UE4M4?
Legacy ROCmFP4 FAST stores 32 four-bit Codebook10 weights plus one UE4M3 scale in 17 bytes. UE4M4 uses the same 17-byte / 4.25-BPW geometry and the same Codebook10 values, but repurposes the unused scale high bit as a fourth scale mantissa bit:
Property
Legacy ROCmFP4 FAST
UE4M4 FAST
Values per block
32
32
Packed weight bytes
16
16
Scale bytes
1
1
Scale encoding
UE4M3
unsigned E4M4
Total block bytes
17
17
Nominal BPW
4.25
4.25
UE4M4's scale candidate set contains every incumbent UE4M3 scale. Under the same exhaustive reconstruction objective, it can select the legacy answer whenever that is best while gaining finer scale placement when useful.
What is in ChadrockV3?
This is the public Q4_0_ROCMFP4_UE4M4_STRIX preset, also exposed as Q4_0_ROCMFP4_UE4M4_STRIX_FAST:
Tensor allocation
Type
431 major model tensors, including routed experts, shared experts, output, and attention K/V
Q4_0_ROCMFP4_UE4M4_FAST
token_embd.weight
Q6_K
Norms, router/control tensors, and other protected tensors
ChadrockV3 is a quantized derivative, not a fine-tune. It does not change Ornith's trained parameters, architecture, tokenizer, or license.
Reconstruction result
The research gate selected quantization state using calibration importance and evaluated the frozen reconstruction with held-out activation importance. The Ornith row covered 10 real BF16 tensors across 7 tensor classes and 3,840 deterministic sampled groups, with no uniform fallbacks.
Geometry
Candidate
Held-out weighted error / legacy incumbent
Error reduction
FAST, 4.25 BPW
UE4M4
0.859755
14.024%
Dual-scale research control, 4.50 BPW
UE4M4
0.819538
18.046%
The released GGUF uses the FAST 4.25-BPW UE4M4 block, not the 4.50-BPW dual-scale research control.
This establishes substantially lower sampled reconstruction error at like-for-like block storage. It does not establish complete-model perplexity, downstream benchmark quality, cross-family universality, or a statistically characterized user-preference gain.
gfx1151 throughput
Measured on an AMD Radeon 8060S / Strix Halo (gfx1151) with ROCm, -ngl 999, flash attention on, --no-mmap, batch 2048, ubatch 512, target-only/no MTP, and three repetitions per geometry.
Artifact
PP512 tok/s
PP2048 tok/s
TG128 tok/s
BF16 teacher
545.19 ± 4.43
534.35 ± 2.52
21.6127 ± 0.0028
Legacy ROCmFP4 STRIX_LEAN
1233.69 ± 9.08
1166.71 ± 7.57
67.7984 ± 0.4843
ChadrockV3 UE4M4 STRIX FAST
1258.82 ± 13.58
1213.08 ± 36.70
67.9382 ± 0.0242
ChadrockV3 vs legacy, mean
+2.04%
+3.97%
+0.21%
Median deltas were +2.09% PP512, +4.15% PP2048, and -0.075% TG128. Decode should therefore be read as performance parity, not as a robust speedup. The PP gains are promising, but PP2048 had 3.03% run-to-run variation and should be replicated on more systems.
The product-level control is near-size-matched, not format-isolated: legacy STRIX_LEAN used Q5_K embeddings and legacy dual-scale attention K/V, while ChadrockV3 uses Q6_K embeddings and UE4M4 FAST attention K/V. ChadrockV3's measured model buffer was 0.275% larger.
Required runtime
Stock llama.cpp and legacy ROCmFPX builds do not know GGML type 108 / ftype 121. Use the public UE4M4 FAST fork:
--no-mmap is required for the validated Strix Halo shared-memory configuration. Tune context and batch geometry for your host. The model's native context metadata is 262,144 tokens; that does not guarantee every context size fits every machine.
Compatibility and limitations
Validated backend: ROCm on gfx1151 / AMD Radeon 8060S.
The public fork includes the UE4M4 producer plus CPU reference/AVX2, generic HIP/CUDA dequantization, GET_ROWS, MMVQ, MMQ, and routed MUL_MAT_ID consumers.
Vulkan and UE4M4 KV-cache registration are intentionally unsupported.
This GGUF has no MTP layer and this release did not test speculative decoding.
This repository does not include a multimodal projector; treat this artifact as text-only.
Full comparative PPL, downstream coding benchmarks, long-context soak, concurrent server load, and broad GPU portability remain unqualified.
The inherited GGUF metadata reports general.name=Ornith-1.0-9B; descriptor/parameter inspection identifies the artifact as the 34.66B-parameter qwen35moe model. The metadata discrepancy is preserved rather than silently rewritten.
Quantized locally from the complete two-shard BF16 GGUF bundle
Hero media: generated with MiniMax H3 FL2VA using an original electro-house soundtrack; the exact Chad Rock chant and caption were added locally. No Party Rock Anthem recording or melody is included.
License: MIT, inherited from Ornith
Machine-readable publication evidence is included in evidence/chadrockv3-evidence.json, and artifact hashes are in SHA256SUMS.txt.
Naming
ChadrockV3 is the release name for this ROCmFP4 UE4M4 STRIX FAST quantization recipe. It is not an upstream Ornith version number and does not imply affiliation with the Ornith or Unsloth teams.
Citation and credit
Ornith is the work of the Ornith/DeepReinforce team. Please cite the original model and follow its model card for intended use:
bibtex
1@misc{ornith-35b,
2 title = {{Ornith-1.0-35B}: Agentic Coding, Open to All},
3 url = {https://deep-reinforce.com/ornith_1_0.html},
4 author = {{DeepReinforce Team}},
5 year = {2026}
6}
UE4M4 integration and ChadrockV3 packaging: hstolte / hstolte11-collab.