Benchmark-aligned validation set for QuaDMix proxy model, designed to improve signal-to-noise ratio compared to the original CORE-22tasks (v1).
Analysis of the 21 CORE benchmark tasks revealed that most tasks have very low answer-to-context ratios (average 5.7%), causing the 1M-parameter proxy model to produce near-random predictions (perplexity ~4800). This resulted in poor LightGBM generalization (val_r2 = 0.179 vs 0.825 for… See the full description on the dataset page:
https://huggingface.co/datasets/liujin99/quadmix-core-bmk-v2.