Ground-truth drafter/premise data lifted from existing Mathlib proofs by
Training/ExtractData.lean. This is the raw pool (429303 draft rows,
664445 premise rows from: Mathlib). Every row is a real have / closed
subgoal taken from a checked proof, so there are no success/used flags to filter
on. Columns follow the uniform schema shared with the rollout dataset, so the two
pools concatenate cleanly.