Test-only Indic rich-transcription benchmark derived from google/fleurs. Each reference transcript is regenerated with grammatical punctuation, formatted numerals, and Indic-script orthographic conventions through an LLM curation pipeline whose prompts were iteratively refined against native-speaker review.
Released alongside the SCRIBE evaluation framework in SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR (accepted at… See the full description on the dataset page:
https://huggingface.co/datasets/adalat-ai/fleurs-ro.