Complete training data for SYZ-Alpha/qwen3.5-0.8b-strict-racket-emitter, a Qwen3.5-0.8B model trained to emit exactly one executable Racket function and no surrounding chatter.
The repository began as a 120-row reviewer sample. The full/ directory now contains every persisted dataset used across the six training stages:
34,421 SFT rows: gold, MultiPL-T breadth, Stage 3 and Stage 4 teacher addenda, and oracle-gated on-policy STaR rows.
6,142… See the full description on the dataset page:
https://huggingface.co/datasets/SYZ-Alpha/Racket-Sample.