You want to fine-tune a small model. There are dozens of base models under 2B parameters. Which one should you invest training time into?
SmolEval answers that question. One command, three metrics — coherence, relevance, diversity — measured on text completion tasks that base models can actually do. No chat templates. No instruction tuning required. Just raw capacity.
BASE CAPACITY TEST (v2.0, 3-run average)
Model… See the full description on the dataset page:
https://huggingface.co/datasets/sifat-febo/smoleval.