I switch source to Q8_0 instead of f16/bf16.
As much as Llama3 models are dense, there is still no notable difference between Q8_0 and fp16 (perplexity is equivalent, or even a little bit lower (-0.001 - -0.002ppl) compared to 16 bpw due to some favorable rounding I guess).