Paper: MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models
Code:
https://github.com/l3cube-pune/MarathiNLP
The L3Cube-MahaParaphrase Dataset is a Marathi paraphrase detection corpus.It is a high-quality, human-annotated corpus specifically designed for Marathi, a low-resource Indic language. It contains 8,000 sentence pairs labeled as either Paraphrase (P) or Non-paraphrase (NP). This dataset is useful for… See the full description on the dataset page:
https://huggingface.co/datasets/l3cube-pune/MahaParaphrase.