Dataset Description:
This dataset is a large-scale collection of Sanskrit Non-STEM textbook data, containing 147 books and 8.29 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and classical knowledge learning in Sanskrit.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page:
https://huggingface.co/datasets/InfoBayAI/Sanskrit-Non-STEM-Textbook-Dataset.