Deduplicated benchmark corpus of 31,591 context snippets (one snippet per single-file source). Use this split for honest evaluation.
Each row has a snippet field with a marker at the boundary position.
Full training corpus: ganga4364/tibetan-outline-boundary-snippets-full
CRF full model: ganga4364/tibetan-outline-boundary-crf-full
CRF unbiased model: ganga4364/tibetan-outline-boundary-crf-unbiased… See the full description on the dataset page:
https://huggingface.co/datasets/openpecha/tibetan-outline-boundary-snippets-unique.