Warmstart SFT dataset for the LoRAcle pipeline (post-pretrain SFT stage).
Built from the union of ceselder/loracle-ia-warmstart and ceselder/loracle-ia-RL
(after excluding the 20-org ceselder/ia-backdoor-trigger-inversion-heldout fair-eval set).
Random 75/25 split of the 883 trainable LoRAs (seed=42):
warmstart_v5 = 75% (662 LoRAs) — this dataset, all rows / varied phrasings
25% (221 LoRAs) held back from warmstart, used for the… See the full description on the dataset page:
https://huggingface.co/datasets/ceselder/loracle-ia-warmstart-v5.