This dataset provides a cleaned and genre-annotated corpus of contemporary Bosnian,
designed for quantitative analysis of language entropy, “language energy”,
and modern NLP tasks.
The canonical release of this corpus is published on Zenodo:
DOI: 10.5281/zenodo.17757098
The corpus is built from three publicly available resources released via the CLARIN.SI repository:
Sarajevo Corpus of SMS Messages in Bosnian 1.1
Bosnian… See the full description on the dataset page:
https://huggingface.co/datasets/hyper-efficient-system-llc/bosnian-corpus-v1.