A synthetic instruction-following dataset of ~984k cleaned, deduplicated conversations generated using Ministral-3B (Apache 2.0 licence) across 9 diverse categories.
Designed as a large-scale pre-training / continual pre-training corpus for small language models, providing high-quality, stylistically varied data that avoids the GPT-4 stylistic bias common in datasets such as OpenHermes and SlimOrca.
This is the larger companion to… See the full description on the dataset page:
https://huggingface.co/datasets/RexiaAI/rexia-synthetic-pretrain-1m.