🌍 TL;DR: MathMist introduces a 21K-sample multilingual benchmark spanning seven languages that enables code-switch CoT and perturbation reasoning analysis in mathematical word problems, revealing how model scale, alignment, and multilingual pretraining jointly shape reasoning performance.
Abstract: Mathematical reasoning remains one of the most challenging domains for large… See the full description on the dataset page:
https://huggingface.co/datasets/mahbubhimel/MathMist.