This repository includes a Tamil-translated versions of the Alpaca dataset and a subset of OpenOrca dataset.
This dataset is part of the release of Tamil LLaMA family of models – an important step in advancing LLMs for the Tamil language. To dive deep into the development and capabilities of this model, please read the research paper and the introductory blog post (WIP) that outlines our journey and the model's potential impact.
GitHub Repository:… See the full description on the dataset page:
https://huggingface.co/datasets/abhinand/tamil-alpaca-orca.