mSCORE is a multilingual, skill-annotated benchmark for evaluating
commonsense reasoning capabilities of large language models. The
benchmark features:
Two sub-benchmarks: mSCoRe-G (General) across 5 languages (en, de,
fr, zh, ja) and mSCoRe-S (Social/Cultural) with tiktok and reddit
splits.
Four complexity levels (L0–L3) created through a structured
context-expansion and implicitation pipeline, progressively hiding… See the full description on the dataset page:
https://huggingface.co/datasets/ngotrnghia1811/mSCORE.