SLM-Bench is a comprehensive, fully programmatic benchmark designed specifically for evaluating sub-10M parameter language models. It covers six core capability areas with 500 questions each (3,000 total), providing granular insights into model strengths and weaknesses.
arithmetic
500
Math problems with plausible distractors (off-by-one, sign errors)… See the full description on the dataset page:
https://huggingface.co/datasets/liodon-ai/slm-bench.