A quality-gated instruction dataset for reasoning, code, and post-training.
CHIMERA Curated 68K is a 68,101-sample instruction-following dataset built for SFT and post-training runs where data quality matters more than raw scale.
The dataset is focused on high-signal reasoning and code examples, with a smaller conversational component included for broader instruction-following behavior. Every admitted sample passes through a quality-gated curation process… See the full description on the dataset page:
https://huggingface.co/datasets/DJLougen/chimera-curated-68k.