Scoring outputs for a small-scale, reviewer-defensible pilot that probes where
multilingual LMs systematically fail on minimal pairs, and whether those
failures reflect grammatical competence or a surface confound (sentence
length / subword tokenisation). Harness:
github.com/suchirsalhan/xblimps.
Published BLiMP-family benchmarks, one common scoring/confound/sampling pipeline,
10 pairs per paradigm (the… See the full description on the dataset page:
https://huggingface.co/datasets/xBLiMPs/xblimps-pilot-results.