Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time… See the full description on the dataset page:
https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-qwen-qwen3-next-80b-a3b-inst-a090ce60.