COILD-MT-Corpus is a collection of sentence-aligned parallel text corpora across 20 language pairs, covering Hindi- and Tamil-centric pairings with other Indian languages. Each pair is stored as two line-aligned plain-text files (one sentence per line, same line count on both sides).
Folder
Language pair
Sentence… See the full description on the dataset page:
https://huggingface.co/datasets/ainlpml-iitp/COILD-MT-Corpus.