This dataset is designed to tackle the core weaknesses of today's large language models when it comes to processing long documents and performing complex reasoning. It consists of 7,500 high-quality training examples across three languages—Chinese, English, and Korean. Each instance is built around a long-text passage and includes questions that require synthesizing information across paragraphs and documents, while… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/Long-Context-Reasoning-Data-Sample.