A sequence-labeling dataset built for fine-tuning BERT-style models (e.g., latin-bert) on Inverse Text Normalization (ITN) — restoring capitalization and punctuation on raw, lowercased Latin text (such as njand/wav2vec2-xls-r-latin ASR outputs).
Compiled from 2,141 files in the CLTK Latin Library and augmented with transcripts from the njand/llpsi-speech-dataset (currently private). Cleaned and transformed through a specialized classical Latin… See the full description on the dataset page:
https://huggingface.co/datasets/njand/latin-asr-post-processing-dataset.