Local RTX 3090 Ti evaluation of q8/q8 KV cache against TurboQuant asymmetric V-cache modes from TheTom/llama-cpp-turboquant.
TurboQuant V-cache behaves like a memory-capacity trade: it substantially improves context headroom, but on this Qwopus 27B MTP setup it is not free. The longer 8K/16K follow-up confirms the speed cost is real and stable when decode length is long enough to measure… See the full description on the dataset page:
https://huggingface.co/datasets/sjakek/qwopus-turboquant-kv-eval.