Georgian Common Voice — telephony-augmented STT dataset
Built from Mozilla Common Voice (ka). The train split contains a clean and a
telephony-degraded copy of each clip (bandpass, u-law codec hops, noise,
packet-loss dropouts). Validation and test splits are clean only.
Text is normalized: numbers verbalized (Georgian vigesimal system),
punctuation and non-Mkhedruli characters stripped.