This dataset is designed for fine-tuning the SanskritBERT model on Part-of-Speech (POS) tagging tasks. It contains sentences tokenized and annotated with POS tags.
The dataset provides two variants of annotations:
Rich: Contains detailed morphological information or fine-grained POS tags, suitable for tasks requiring deep linguistic analysis.
Simple: Contains simplified POS tags (likely Universal POS tags), suitable for general-purpose tagging and… See the full description on the dataset page:
https://huggingface.co/datasets/tanuj437/sanskrit-pos-tagged-corpus.