A length-filtered subset of the no_think split of
openbmb/UltraData-SFT-2605,
containing conversations whose token length falls in the 8k–32k range.
This is the medium-length tier intended for standard long-context SFT.
Two companion tiers were produced from the same source:
fxmeng/UltraData-SFT-2605-no-think-32k-200k… See the full description on the dataset page:
https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-8k-32k.