17 million samples · 6 source corpora · token-level sequence labeling
Accepted at the EACL 2026 SilkRoad NLP Workshop
Punctuation restoration is essential for improving the readability and downstream utility of automatic speech recognition (ASR) outputs, yet remains underexplored for Persian despite its importance. We introduce PersianPunc, a large-scale, high-quality dataset of 17… See the full description on the dataset page:
https://huggingface.co/datasets/MohammadJRanjbar/PersianPunc.