A synthetic OCR benchmark for written Cantonese (廣東話).
粵語(廣東話)OCR 合成評測資料集。
English · 繁體中文
4,930 synthetic text-line images testing whether vision-language models can read
the characters used in written Cantonese. Everyday writing in Hong Kong… See the full description on the dataset page:
https://huggingface.co/datasets/trickster-2005/cantonese-benchmark-synth.