Everything needed to replicate the data generation, the biasing experiments and the
contrastive-head training, in one folder.
repos/ source code (full git history included)
speech-data-gen/ data + audio generation pipeline
speech-model-evaluation/ speechllm evaluation harness
vllm-scripts/ vLLM / k8s serving scripts
A self-contained browser… See the full description on the dataset page:
https://huggingface.co/datasets/chaurAr/Cross-lingual-entity-M.