Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required
... and final boxed answer after reasoning
exact text/problem/response dedupe and near problem… See the full description on the dataset page:
https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.