Views
No views yet
<no_target>) when the target is absent.asr_backbone.pt — recognizer + activity-detection backbone: a frozen ECAPA-TDNN speaker embedding modulates a WavLM-Base+ encoder via FiLM, feeding two language-specific CTC heads (English/Latin, Kazakh/Cyrillic) and a frame-level VAD head.presence_gate.pt — utterance-level target-presence gate (enrollment–mixture matching + speaker-conditioned attention + attentive-statistics pooling), applied on top of the frozen backbone.config.json — backbone configuration (vocab sizes, hyperparameters).| Test set | Raw WER | Gated WER | Detection BAcc |
|---|---|---|---|
| English (Libri3Mix-100h) | 29.35 | 36.92 | 80.04 |
| Kazakh (Kazakh3Mix-100h) | 43.47 | 50.71 | 86.78 |
1@article{persona_asr,
2 title = {Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh--English Overlapping Speech},
3 year = {2026}
4}