This dataset is a copy of PowerInfer/QWQ-LONGCOT-500K.
This repository contains approximately 500,000 instances of responses generated using QwQ-32B-Preview language model. The dataset combines prompts from multiple high-quality sources to create diverse and comprehensive training data.
The dataset is available under the Apache 2.0 license.
Over 75% of the responses exceed 8,000 tokens in length. The majority of prompts were carefully created using persona-based methods to create challenging… See the full description on the dataset page:
https://huggingface.co/datasets/huihui-ai/QWQ-LONGCOT-500K.