Text corpora for building token-level n-gram language models for
Lithuanian ASR beam search decoding. Used with
sliderforthewin/parakeet-tdt-lt.
europarl.sentences.txt.zst
634k
EU… See the full description on the dataset page:
https://huggingface.co/datasets/sliderforthewin/lt-asr-lm-corpora.