Views
No views yet
⛔ THIS BUILD DOES NOT FIT ON A 128 GB STRIX HALO
llama.cppreports 113.03 GiB addressable on a Ryzen AI MAX+ 395. These 8-bit builds are 114.38 GiB and 116.15 GiB. Attempting-ngl 999hard-wedges the machine — we did it twice: a KFD SVM D-state livelock (svm_range_cpu_invalidate_pagetables) that survives a GPU reset and needs a power cycle.-fit offdoes not save you; it only stops llama.cpp from shrinking the model, so it allocates until the driver dies.On a single 128 GB Strix Halo, use the 4-bit build (63.07 GiB, 37.86 tok/s). These 8-bit builds are for machines with more memory, or for CPU / partial-offload inference.
| File | Mistral-Small-4-119B-2603-Q8_0_ROCMFPX_AGENT.gguf |
| Size | 116.15 GiB |
| BPW | 8.38 |
| ftype | Q8_0_ROCMFPX_AGENT (115) |
| Tensors | 579 |
--output-tensor-type q8_0 --token-embedding-type q8_0 --tensor-type shexp=q8_0.
The shexp override matched 108 shared-expert tensors (confirmed in the dry-run receipt —
a --tensor-type pattern that matches nothing is a silent no-op, so we check the count).Q8_0_ROCMFPX (111) / Q8_0_ROCMFPX_AGENT (115) exist only in
charlie12345/ROCmFPX. Stock llama.cpp reports
invalid ggml type 103. Ignore the auto-generated "Use this model" commands above.| variant | ftype | size | bpw | GPU on 128 GB Strix Halo | decode |
|---|---|---|---|---|---|
| 4-bit COHERENT | 102 | 63.07 GiB | 4.55 | ✅ fits | 37.86 tok/s |
| 8-bit AGENT | 115 | 116.15 GiB | 8.39 | ⛔ does not fit | not measurable on this box |
| 8-bit plain | 111 | 114.38 GiB | 8.26 | ⛔ does not fit | not measurable on this box |
-ngl 0): 17×23 ⇒ ✅ 391 · capital of Japan ⇒ ✅ Tokyo.
Loaded in 256 s from disk.