This dataset was created by preprocessing the
allenai/tulu-3-sft-mixture
dataset for single-turn supervised fine-tuning.
The preprocessing keeps English user -> assistant examples from the selected
Tulu sources, applies length filtering with the official
Qwen/Qwen3.5-4B-Base chat template, and removes exact and near duplicates.
The resulting train split contains 151,292 examples with a maximum sequence
length of 7,168 tokens.
Each row contains the… See the full description on the dataset page:
https://huggingface.co/datasets/HwanChang0106/tulu_sft_mixture_preprocessed.