TransCorpus-bio-hi is a large-scale, parallel biomedical corpus consisting of Hindi synthetic translations of PubMed abstracts. This dataset was created using the TransCorpus framework and is designed to enable high-quality Hindi biomedical language modeling and downstream NLP research.
Source: PubMed abstracts (English)
Target: Hindi (synthetic, machine-translated)
Translation Model: M2M-100 (1.2B) using TransCorpus Toolkit
Size: 22… See the full description on the dataset page:
https://huggingface.co/datasets/jknafou/TransCorpus-bio-hi.