This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
This dataset contains English-to-Tamil parallel sentence pairs designed for machine translation and text generation tasks. The samples cover diverse domains including news, government schemes, religious texts, and casual conversation. It is a subset of the larger Samanantar corpus, featuring source English prompts and corresponding Tamil completions.… See the full description on the dataset page:
https://huggingface.co/datasets/sharanR2005/adaption-samanantar-en-ta.