Input dataset for the logprob_pivots track: a fixed 50-problems-per-language MGSM subset so all runs score the same problems.
provenance
python -m scripts.logprob_pivots.build_mgsm_subset (seed 42, n_per_lang 50, src/logprob_pivots/config.py)
schema
CSV: one row per problem x language with… See the full description on the dataset page:
https://huggingface.co/datasets/multicot/2026-01-01-mgsm-subset-en-es-fr-de.