This is the finalized V4 Master Collection for the Zenyx project, expanding upon the previous V3. This version focuses on high-reasoning, code, and math capabilities through massive distillation and thinking-chain integration.
Total Samples: 101,523
Branding: Rebranded to Zenyx / Zenyx Lab.
Filtering: Strict English, Math, and Code filter applied. Chinese and non-standard characters removed.
Format:… See the full description on the dataset page:
https://huggingface.co/datasets/Arko007/zenyx-v3-synthetic-sft-v4.