Views
No views yet

|
Q8_0 · Recommended Best balance of quality, speed, and memory. Fastest single-request result in our tests. 323 MiB Download Q8_0 |
F16 · Highest fidelity Closest practical copy of the source adapter. Best measured DSpark and batched throughput. 608 MiB Download F16 |
NVFP4 · Experimental Smallest build, but lower fidelity and a custom CUDA backend are required. Not for stock llama.cpp. 171 MiB Download NVFP4 |
This repository contains LoRA adapters only. Download a compatible Qwen3.6-27B base model separately. Testing usedTernary-Bonsai-27B-Q2_0.gguf.
1hf download prism-ml/Ternary-Bonsai-27B-gguf \
2 Ternary-Bonsai-27B-Q2_0.gguf --local-dir .
3
4hf download ajh-code/Ternary-Bonsai-27B-Grug-LoRA-GGUF \
5 Grug-LoRA-r64-Q8_0.gguf --local-dir .
6
7llama-server \
8 -m Ternary-Bonsai-27B-Q2_0.gguf \
9 --lora-scaled Grug-LoRA-r64-Q8_0.gguf:1.5 \
10 -ngl 99 -c 327681.5 is recommended for Ternary Bonsai. It compensates for the changed
ternary base and produced the best completion reliability and token efficiency
in our calibration. Scale 1.0 matches the source-adapter baseline, but was
too weak on this base.| Adapter | Single stream | DSpark, high acceptance | Batch 32 aggregate | Peak VRAM |
|---|---|---|---|---|
| F16 | 65.1 tok/s | 158.8 tok/s | 299.7 tok/s | 8,218 MiB |
| Q8_0 | 73.9 tok/s | 152.4 tok/s | 290.4 tok/s | 7,868 MiB |
| NVFP4 | 71.6 tok/s | 154.6 tok/s | 285.3 tok/s | 7,716 MiB |
llama-server, deterministic behavior checks, and two preallocated
134,144-token KVarN K4/V4 slots. Full methods and raw results are maintained
in the companion research project.