✨ Note: For all FineInstructions resources please visit:
https://huggingface.co/fineinstructions
This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline.
The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). See our paper for more details.
Each .parquet file in the data folder has a corresponding judge-*.json file that… See the full description on the dataset page:
https://huggingface.co/datasets/Bas95/fineinstructions_nemotron.