This release contains a cleaned full training set and a frozen evaluation set.
Audio is embedded in Parquet and can be streamed directly by the Hugging Face
datasets library.
The train split contains 3,209,619 cleaned base rows.
No external code-switch corpus is included in this release. The frozen eval audio is physically separated from… See the full description on the dataset page:
https://huggingface.co/datasets/WTForbes/nan_new.