qcache-q8_0.v1.gguf — GGML Q8_0 tensors for the model's linear projections,
stored in a GGUF container keyed by the original checkpoint tensor paths
(no metadata KVs). Produced by in-memory quantization of the bf16
checkpoint with a candle-based
loader.aux-tensors.safetensors — everything the cache does not carry:
embeddings, norms and other non-projection tensors, in their original
dtypes.