Supplementary SFT rows for a pre-1930 conversational model. 7,338 rows across five
routes. Kept deliberately separate from the graded synthetic dataset, whose rows are
answers lifted verbatim from period prose and scored by a judge; these are constructed,
ungraded, and would muddy that provenance guarantee if mixed in.
A model pretrained on pre-1930 books and finetuned on question-and-answer pairs behaves
correctly right up… See the full description on the dataset page:
https://huggingface.co/datasets/zachnorton03/vintage-sft-robustness.