SFT warmup dataset for baseline GRPO condition. 1200 Countdown problems (5 numbers, +/-/*) solved by Qwen3-1.7B with 32k token generation. Contains full untruncated reasoning traces. No confidence tags.
messages
List({'content': Value('string'), 'role': Value('string')})
Chat-format conversation (system, user, assistant). System sets task format, user provides… See the full description on the dataset page:
https://huggingface.co/datasets/reasoning-degeneration-dev/wmc-sft-baseline-v2.