Agriculture SLM Corpus v3
Pre-training corpus for a 300M parameter India-focused agriculture SLM.
File
Description
Docs
Tokens
corpus_india_downsampled.jsonl
India-focused downsampled base corpus (PubMed, EuropePMC, Wikipedia, extension, ICAR, etc.)
~201K
~1,244M
krishikosh.jsonl
ICAR KrishiKosh theses, articles, books, reports (DSpace 7 harvest)
~55K
~1,385M
{
"text": "...",
"source": "krishikosh | ..."… See the full description on the dataset page:
https://huggingface.co/datasets/AnmolNimmala0/agri-slm-corpus-3.