This the 200k subset of the 'unfiltered' version of the Tulu v2 SFT mixture, created by collating the original Tulu 2 sources and avoiding downsampling.
This was used for the 200k-size experiments.
FLAN (Apache 2.0, we only sample 961,322 samples along with 398,439 CoT samples from the full set for this data pool)
Open Assistant 1 (Apache 2.0)
ShareGPT (Apache 2.0 listed, no official repo… See the full description on the dataset page:
https://huggingface.co/datasets/hamishivi/200k-tulu-2-unbalanced.