MLP-based expert prediction models trained to prefetch MoE experts before they're needed. This is an informative negative result: despite achieving 86.65% hit@8 prediction accuracy, expert prefetch provides zero speedup on consumer hardware because CPU GEMV for small experts (430 KB) takes <50μs — less than 45% of the compute budget.
expert_predictor/
35 MB
Base: same-layer prediction… See the full description on the dataset page:
https://huggingface.co/datasets/Kevletesteur/chimere-expert-predictor.