Evaluation of SOTA models on the LIMO dataset (v1).
LIMO is a small-scale dataset (817 Q&A) that was originally used for SFT LLMs to acquire reasoning capabilities.
This repo contains the answers provided by several recent models evaluated on the LIMO dataset, and the objective of it is to show that maybe, right now,
we are reaching a point where, nonreasoning models are performing well enough to make this dataset less appealing.
In the original LIMO paper, the LIMO dataset was used to… See the full description on the dataset page:
https://huggingface.co/datasets/fedric95/LIMO-Traces.