Maskara Indian PII Dataset (Phase 2)
Synthetic Indian PII dataset for training the Maskara NER model. This version has been expanded and diversified for Phase 2 training.
Split
Rows
Purpose
train
~250,000
Model training
template_disjoint_eval
15,000
Generalization evaluation: entire template families held out
real_world_eval
~2,600
Manually curated real-world Indian text evaluation
Column
Type
Description… See the full description on the dataset page:
https://huggingface.co/datasets/somukandula/maskara-indian-pii-200k.