Fine-tuned ECAPA-TDNN for Vietnamese Speaker Verification
This model is a fine-tuned ECAPA-TDNN speaker embedding model based on
speechbrain/spkrec-ecapa-voxceleb.
Training
- Fine-tuning dataset: Vietnam-Celeb
- Selected checkpoint: Epoch 10
- Embedding dimension: 192
- Sample rate: 16 kHz
Verification Pipeline
Audio
→ CRDNN VAD
→ ECAPA-TDNN embedding
→ L2 normalization
→ 5-recording enrollment
→ mean speaker centroid
→ cosine similarity
→ threshold decision
Final DEV-calibrated threshold:
0.1566
Speaker-disjoint Evaluation
All-impostor evaluation protocol:
| Model | DEV EER | TEST EER | TEST FAR | TEST FRR |
|---|
| Pretrained ECAPA | 13.30% | 11.91% | 9.46% | 13.47% |
| Fine-tuned Epoch 10 | 9.98% | 8.42% | 9.00% | 7.88% |
TEST evaluation:
- 50 unseen speakers
- 5 enrollment recordings per speaker
- 698 genuine trials
- 34,202 impostor trials
- No speaker overlap between train/DEV/TEST
The TEST threshold was not tuned on TEST.
The final threshold was calibrated only on DEV.
Base Model
SpeechBrain:
speechbrain/spkrec-ecapa-voxceleb
Limitations
This model is intended for an academic prototype of a speaker-verification
system. The reported FAR is not low enough to claim production-grade
security.
Performance may vary with microphone quality, background noise, language,
recording duration, and speaker characteristics.
Files
ecapa_vietnamceleb_epoch10.pt: fine-tuned ECAPA embedding checkpoint
config.json: deployment configuration and verification threshold
all_impostor_metrics.json: detailed evaluation metrics
pretrained_vs_finetuned.csv: baseline comparison