This is a clustered version of the allenai/tulu-3-sft-mixture dataset, where each example is assigned to one of the original Tülu-3 paper categories based on its source field.
Each source dataset was mapped to exactly one category based on the definitions in the Tülu-3 paper. No heuristic text-based clustering was used.
The goal is to make it easy to train per-category experts or do category-aware sampling for instruction tuning and… See the full description on the dataset page:
https://huggingface.co/datasets/r-three/tulu3-sft-og-clustering.