SYNTHETIC-1 is a reasoning dataset obtained from Deepseek-R1, generated with crowdsourced compute and annotated with diverse verifiers such as LLM judges or symbolic mathematics verifiers. This is the SFT version of the dataset - the raw data and preference dataset can be found in our 🤗 SYNTHETIC-1 Collection.
The dataset consists of the following tasks and verifiers that were implemented in our library… See the full description on the dataset page:
https://huggingface.co/datasets/PrimeIntellect/SYNTHETIC-1-SFT-Data.