Chinese ↔ 通用盲文 (Chinese Braille) parallel training data for the Vision-Braille translation
models. Three corpora ship together in this repo:
Folder
What it is
Rows
Size
cleaned_2345_v3/
Audited + cleaned braille→Chinese training corpus, stages 2→5, with restored 分词连写 word spacing on stage 2
379,152
426 MB
braille_spaced/
Freshly built passage-level corpus, fully word-spaced, across 通用 / 数学 / 医学 / 中医 / 病理学
80,360
208 MB
cleaned_080126_v1/… See the full description on the dataset page:
https://huggingface.co/datasets/Violet-yo/vision-braille-dataset.