Sangraha Tamil (Cleaned & Verified)
Dataset Description
This dataset is a cleaned, processed subset of the AI4Bharat Sangraha (Verified) dataset, specifically targeting the Tamil language. It was prepared for the purpose of Continuous Pre-Training (CPT) of Large Language Models (LLMs) like Llama 3 and Qwen 2.5/3 to improve their performance on Indic languages.