This is a spoken corpus of Hill Mari language assembled by this project.
Total amount of tokens: 63522. It mostly contains texts gathered in a Hill Mari expedition under Egor Kashkin.
Please make sure to cite them:
@online{hill_mari_msu_corpus,
author = {Айгуль Закирова and Анастасия Гарейшина and Анастасия Сибирёва and Анастасия Сиротина and Анита Соловьёва and Анна Бочкова and Вадим Дьячков and Владимир Иванов and Дарья Белова and Дарья Мордашова and… See the full description on the dataset page:
https://huggingface.co/datasets/OneAdder/hill-mari-msu-spoken-corpus.