A standardized evaluation benchmark designed specifically for mobile and edge-deployed language models.
Existing benchmarks (MMLU, HumanEval, GSM8K) test what large models can do on servers. MobileBench tests what small models can do on phones — the tasks users actually perform:
Summarization — The #1 on-device task (messages, emails, notifications)
Classification — Spam detection, sentiment, intent… See the full description on the dataset page:
https://huggingface.co/datasets/dispatchAI/MobileBench.