Domain-specific LLM training
Legal question answering
Policy reasoning & case retrieval
Agentic systems for legal workflow automation
This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including:
English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others.
The corpus is structured and processed to be… See the full description on the dataset page:
https://huggingface.co/datasets/antonhome/indian-legal-supervised-fine-tuning-data.