The Chinese Classification Dataset is designed for classifying sentences in different forms of Chinese text. Each entry consists of a Chinese sentence paired with a label indicating its language variant.
Structure
Rows: Each row contains a single Chinese sentence.
Labels: Comma-separated strings indicating one or more of these languages:
zh: Simplified Chinese
zht: Traditional Chinese
yue: Cantonese
Example… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/chinese-classification.