Full model responses on the MATH held-out test split for two models, to support
behavioural auditing. All transcripts are from a single default condition (a standard
step-by-step solve prompt; no special system prompt or prefix).
This is one of a pair of sets derived from a common transcript pool. Each set contains the
same trusted model and one model under investigation; the sets do not disclose how the
two investigated models relate… See the full description on the dataset page:
https://huggingface.co/datasets/darklord1611/math-eval-transcripts-b.