Text: 9000 records in text/cs_text_dataset.jsonl
Audio: 623 clips in audio/
Splits: manifests/splits.json
Metadata: manifests/DATACARD.json, manifests/hf_dataset_info.json, manifests/summary.json
Text pipeline: T0 normalize -> T1 rule constraints -> T2 copy-switch bootstrap -> T4 filters
Audio pipeline: A0 model setup ->… See the full description on the dataset page:
https://huggingface.co/datasets/nnoukastephen/CodeSwitch-SW-LIN-FRA-synthetic-audio.