Apache-2.0 training and evaluation assets for the four-stage Kakeya OProver LoRA journey.
sft-90: 63 train / 9 validation / 18 holdout
sft-1000: 700 train / 100 validation / 200 holdout
dpo-2000: 1,400 train / 200 validation / 400 holdout preference pairs
repair-2800: 2,800 train / 400 validation (train: 2,100 repair + 700 rehearsal; validation: 300 + 100)
strict-evaluation-22: 22 public-safe task identifiers and… See the full description on the dataset page:
https://huggingface.co/datasets/FluffyAIcode/Kakeya-OProver-Training.