Two monolingual corpora built for a pair of ~25M-parameter decoder-only
Transformers. Hindi is the higher-resource language, Nepali the lower-resource
one. Both are written in Devanagari (U+0900-U+097F), so script cannot be used
to tell them apart --- separating them is the central technical problem this
dataset solves rather than assumes.
language
documents
characters
manual (chars)
tokens
manual (tokens)
train
val
test… See the full description on the dataset page:
https://huggingface.co/datasets/meet5568/lma_datasets.