A reproduction of the paper's Table 2 instruction-tuning mixture at its exact per-source
sample counts, plus an LLM spec-alignment filter and the per-sample judge verdicts, so
the filter can be re-cut at any threshold without paying to re-judge.
field
value
experiment
Table 2 instruction-tuning mixture for the Teaching Claude Why replication, filtered for spec misalignment