A curated subset of the AI4Bharat IndicVoices-R dataset created for
Indian-language speech classification experiments on consumer hardware.
The original IndicVoices-R corpus is extremely large (~750 GB). This
dataset intentionally samples a much smaller subset so that researchers
and developers can download, preprocess, and experiment with the data
locally on machines with limited storage and memory.