Synthetic instruction and pre-training style dataset prepared for Hugging Face Hub. The repository contains three train configs with intentionally different supervision styles and non-repeating prompt text across shards.
This dataset has three configs:
Config
Focus
Samples
train-001
Instruction-following and task completion
960
train-002
Transformation, labeling, repair, and ranking
576
nous
Larger mixed supervision corpus
4,200
Loading… See the full description on the dataset page: https://huggingface.co/datasets/Surpem/Nous-Instuct-PT.