This repository contains two configurations of Nepali LLM datasets:
Configurations
1. Scrapy Engine
Description: Contains data collected using a web scraping engine.
Files: [List any specific files or formats]
2. Nepberta
Description: This dataset is derived from the Nepberta project and contains cleaned data specifically related to the project. The dataset contains **cleaned text chunks of size ~50 mb ** of all… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/nepali_llm_datasets.