This dataset contains writings in (possibly) a mixture of Standard Chinese and Cantonese, derived from the NLPTEA 2017 Chinese Spelling Check Shared Task (Fung et al., NLP-TEA 2017).
This dataset is intended for text classification or token classification (span detection) tasks.
Columns:
id: Identifier, in the format of ASTRI0XXX, EVAXXX or ADDXXX.
sentence: A Chinese sentence that may contain spelling mistakes and/or Cantonese colloqialisms.… See the full description on the dataset page:
https://huggingface.co/datasets/Swithord/cantonese-span-detection.