A small, fully-reproducible study on a question everyone assumes they know the answer to:
does fine-tuning a small model on a benchmark make it better?
Short answer: it depends entirely on the training data, not on the fact that you fine-tuned.
The same LoRA recipe, run with two different kinds of data, gives opposite signs.