For more details on talkie, see
their website.
If you're looking to run talkie at home, I'd recommend using Q3_K for 8GB VRAM, Q4_K for 10GB VRAM, Q5_K for 12GB VRAM.
The smaller subvariants will allow for a larger context length at the cost of more degradation.
I was able to run Q3_K_L on my GTX 1070 8GB at 512 context length using Q8 KV cache.
Talkie has large outliers which make quantization difficult. I ran a search to optimize the allocation of precision between tensors for a given model size.
I found that layer 14's FFN down projection in particular causes issues and should be left in a higher precision.
The imatrix was trained on
dell-research-harvard/AmericanStories, using pre-1930 data.
llama.cpp uses a higher precision for activations than the reference PyTorch implementation, resulting in some KL divergence even with Q8_0 and BF16.
This shouldn't affect output quality.