Pre-tokenized version of the allenai/Dolci-Think-SFT-7B dataset, ready for training with OLMo-core.
This dataset was used to train the openeurollm/OLMo-3-7B-Think-SFT checkpoints.
See also: openeurollm/dolci-instruct-sft-tokenized for the instruct (non-thinking) variant.
Total instances
2,268… See the full description on the dataset page:
https://huggingface.co/datasets/openeurollm/dolci-think-sft-tokenized.