Pre-training corpus for a 300M parameter India agriculture domain Small Language Model (SLM).
Total tokens: ~3.45B (GPT-2 tokenizer)
Total documents: 556,765 (Train: 552,203, Test: 4,562)
Language: English only
Domain: Agriculture — India-specific
Shards: 124 train shards + 1 test shard
Each document is labeled with one or more of 29 fine-grained agriculture categories and a… See the full description on the dataset page:
https://huggingface.co/datasets/AnmolNimmala0/agri-slm-india-v1.