A curated 100,000-row SFT mixture dataset for instruction-tuning LLMs.
Logic Core (60%)
PrimeIntellect/SYNTHETIC-2-SFT-verified
60,000
Reasoning, math, coding with verified DeepSeek-R1 traces
Conversational Glue (30%)
Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered
30,000
Multi-turn fluidity, tone alignment
Human Element (10%)
allenai/WildChat-4.8M
10,000
Real-world prompts… See the full description on the dataset page:
https://huggingface.co/datasets/TeamClaude/sft-100k-mixture.