Views
No views yet
[!CAUTION] This quant uses custom ROCmFP4 tensor formats + Laguna architecture support. It does NOT load in stock upstreamllama.cpp. You MUST use the specialized Ciru ROCmFPX Runtime V3.
agent/laguna-s21-runtime-v354f5fe06c74350fb8b6aec21d8749071bc195bdbgfx1151) with 128 GB Unified Memory.Vulkan0 via Mesa RADV driver).laguna-s-2.1-ROCmFP4-StrixKVSpine-v4.gguf (65.4 GB).1git clone --branch agent/laguna-s21-runtime-v3 --depth 1 https://github.com/ciru-ai/ROCmFPX.git
2cd ROCmFPX
3git checkout --detach 54f5fe06c74350fb8b6aec21d8749071bc195bdb1scripts/install-laguna-vulkan-deps.sh --install
2JOBS=8 BUILD_TYPE=Release scripts/build-laguna-strix-vulkan.shHOST=0.0.0.0 PORT=8089 scripts/run-laguna-vulkan-supervised.sh /path/to/laguna-s-2.1-ROCmFP4-StrixKVSpine-v4.ggufhttp://host.docker.internal:8089/v1 or http://172.18.0.1:8089/v1http://127.0.0.1:8089/v1sk-no-key-required (or any string)laguna-s21-rocmfp4-strixkvspine-v4| Metric | Result |
|---|---|
| Generation Speed (Decode) | 33.1 – 38.3 t/s |
| Prompt Eval Speed (Prefill) | 79.3 – 92.7 t/s |
| Time to First Token (TTFT) | ~560 – 618 ms |
| Validated Max Context | 131,072 tokens (131K) |
| Default Thinking | Off |
Q4_0_ROCMFP4_StrixKVSpine-v4 preset)ea1d854a72c47ec8e72c16ea91b8ff3cd5e1620b834df175f683c86f27dc26d6halofpx pull laguna-s21 && halofpx serve -m laguna-s21 (Vulkan0 backend, 131K context)