Every measured row behind the stillwarm project: does saving/restoring a
llama-server conversation's KV cache to disk actually work, when does it beat
recomputing, and what silently breaks it?
Hardware/build (frozen): MacBook Pro, Apple M3 Max, 36 GB unified memory,
macOS 26.5.1; llama.cpp release b9871 (ef2d770…), Release build, Metal;
models pinned by SHA-256 (Llama-3.1-8B-Instruct Q4_K_M, Qwen2.5-7B-Instruct
Q4_K_M, Gemma-3-4B-it Q4_K_M — the… See the full description on the dataset page:
https://huggingface.co/datasets/vimalnakrani/stillwarm-bench-results.