This is a fine-tuned
whisper-large-v3 model for
Latgalian, trained by
AiLab.lv using two general-purpose speech datasets:
As a base model, we used a previously fine-tuned ASR model for
Latvian, and continued to fine-tune it for Latgalian. The fine-tuning was done using the Hugging Face Transformers library.
NB! The MuLaR corpus contains transcriptions that generally do not follow the rules of the standard Latgalian orthography, in contrast to the Latgalian CV corpus.
This work was supported by the EU Recovery and Resilience Facility project
Language Technology Initiative (2.3.1.1.i.0/1/22/I/CFLA/002) in synergy with the State Research Programme project "Diversity of Latvian in Time and Space" (VPP-LETONIKA-2021/4-0003).