Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
At 1,075,000 rows (5.375× the XL variant), this is the large-scale training dataset for serious fine-tuning runs. The full topic bank is sampled densely, providing high repetition for core topics and meaningful coverage of rare ones.
Sharded into 250,000-row JSONL files for easy… See the full description on the dataset page:
https://huggingface.co/datasets/Nix-ai/cat-v3xxl.