Hindi Language Pre-Trained LLM Datasets Overview
Welcome to the Hindi Language Pre-Training Datasets repository! This README provides a comprehensive overview of various pre-training datasets available for Hindi, including essential details such as licenses, sources, and statistical information. These datasets are invaluable resources for training and fine-tuning large language models (LLMs) for a wide range of natural language processing (NLP) tasks.
-Data Overview and Statistics
This README… See the full description on the dataset page:
https://huggingface.co/datasets/Hindi-data-hub/odaigen_hindi_pre_trained_sp.