The dataset contains 3,585 rows.
These paragraphs are extracted from authorized literary works written by Png Iau Khian方耀乾 and contain multiple sentences in Taiwanese Taigi written with Hanji.
The dataset maintains the original literary style and structure, making it useful for training language models, natural language processing (NLP), and Taiwanese literature research.
Number of rows: 3,585 (each representing a paragraph)… See the full description on the dataset page:
https://huggingface.co/datasets/IMA-Taiwan/taigi-literature-pikh.