TL;DR — On a real decoder MoE (allenai/OLMoE-1B-7B-0924, 1024 experts, top-8) with a valid perplexity metric, traffic-aware selective 4-bit quantization works: protecting the highest-traffic (hot) experts at full precision and quantizing the rest Pareto-dominates a same-storage random (linear-mix) allocation at every budget, recovering 46.9% of the uniform-to-bf16 perplexity gap at a ~24% storage premium (hottest 10%) and… See the full description on the dataset page:
https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-12-nvfp4-moe-selective-quant.