This is the deliberately small first experiment: train on verified historical
IOL material, measure on untouched development problems, then run the resulting
checkpoint on the locked 2021–2023 Linguini test and the competition.
final/train.jsonl: 616 chat examples (including three marked repeats so all
records fit complete batches of eight).
final/eval.jsonl: 61 examples from ten held-out pre-2021 problems.
locked/locked_test.jsonl: 32… See the full description on the dataset page:
https://huggingface.co/datasets/ChrisR05/iol-ai-2026-sft-v1-data.