This dataset is designed for instruction fine-tuning of large language models (LLMs), especially for the Qwen3 family, to perform punctuation restoration on Mandarin Chinese text.
It is derived from the AWeirdDev/zh-tw-articles-6k dataset. The context field is processed to create input-output pairs in the Qwen3-style message format.