Dataset Description:
This dataset is a large-scale collection of Telugu STEM textbook data, containing 596 books and 26.00 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Telugu.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for deeper… See the full description on the dataset page:
https://huggingface.co/datasets/InfoBayAI/Telugu-STEM-Textbook-Dataset.