anger, fear, happy, neutral, sad, surprise| Component | Specification |
|---|---|
| Backbone | Wav2Vec2-Base (Facebook / Meta) |
| Pretrained Weights | IEMOCAP Emotion Recognition |
| Parameters | 95.2M total (~856K trainable initially during warmup) |
1Wav2Vec2 Encoder (768-dim output)
2 ↓
3Attention Pooling
4 ↓
5Linear(768 → 512) + LayerNorm + ReLU + Dropout(p=0.3)
6 ↓
7Linear(512 → 256) + LayerNorm + ReLU + Dropout(p=0.3)
8 ↓
9Linear(256 → 128) + LayerNorm + ReLU + Dropout(p=0.3)
10 ↓
11Linear(128 → 6) ────> [Output: 6 Emotions]| Dataset | Description |
|---|---|
| RAVDESS | Ryerson Audio-Visual Database (24 actors, highly structured high-definition clips) |
| IESC | Indian Emotional Speech Corpus (8 speakers × 5 emotions × 10 repetitions) |
| Hindi | Hindi Speech Dataset |
| KAI-Indian | KAI Indian Emotional Speech Corpus |
| URDU-Dataset | Urdu Emotional Speech Dataset |
| UrduSER | Urdu Speech Emotion Recognition Dataset |
| Split | Audio Files | Purpose |
|---|---|---|
| Train (Balanced) | 6,000 | 1,000 samples per class (including Mixup & augmentation) |
| Validation | 1,057 | Original samples only |
| Test | 1,059 | Original samples only |
| Total | 8,116 |
anger : 1,000 samplesfear : 1,000 sampleshappy : 1,000 samplesneutral : 1,000 samplessad : 1,000 samplessurprise : 1,000 samples| Parameter | Value | Purpose |
|---|---|---|
| Batch Size | 8 | Memory-efficient training on single-GPU hardware |
| Gradient Accumulation | 4 | Effective batch size of 32 for stable gradients |
| Learning Rate | 5e-05 | Peak rate suited for transfer learning |
| Max Epochs | 15 | Baseline training duration |
| Freeze Epochs | 2 | Classifier warmup (Wav2Vec2 backbone fully frozen) |
| Early Stopping Patience | 10 epochs | Early termination trigger based on validation loss |
~1e-05.| Epoch | Train Loss | Train Acc | Val Loss | Val Acc | Notes |
|---|---|---|---|---|---|
| 1 | 1.7864 | 22.9% | 1.7539 | 25.9% | Baseline established with random classifier weights |
| 2 | 1.7528 | 26.9% | 1.7508 | 27.5% | Minor classifier adjustments before backbone unfreeze |
| Epoch | Train Loss | Train Acc | Val Loss | Val Acc | Train-Val Gap | Status / Performance |
|---|---|---|---|---|---|---|
| 3 | 1.5474 | 46.4% | 1.3877 | 51.7% | -5.3% (Val higher) | Backbone unfrozen; instant performance leap |
| 4 | 1.4115 | 54.8% | 1.2229 | 62.9% | -8.1% (Val higher) | Rapid model adaptation |
| 5 | 1.2636 | 66.8% | 1.2188 | 62.2% | +4.6% | Slight validation fluctuation |
| 6 | 1.2175 | 70.3% | 1.1253 | 68.7% | +1.6% | Convergence continues |
| 7 | 1.1086 | 76.2% | 1.1493 | 68.5% | +7.7% | No structural metric improvement |
| 8 | 1.0589 | 79.5% | 1.0615 | 72.1% | +7.4% | New Peak |
| 9 | 1.0215 | 81.6% | 1.0815 | 73.4% | +8.2% | New Peak |
| 10 | 1.0136 | 82.8% | 1.0630 | 73.7% | +9.1% | New Peak |
| 11 | 0.9250 | 85.7% | 1.0572 | 74.4% | +11.3% | New Peak |
| 12 | 0.9015 | 88.4% | 1.0972 | 73.2% | +15.2% | Small val dropout |
| 13 | 0.8837 | 90.7% | 1.0632 | 76.2% | +14.5% | New Peak |
| 14 | 0.8524 | 92.3% | 1.0284 | 77.3% | +15.0% | Best Checkpoint (Selected) |
| 15 | 0.7972 | 93.3% | 1.1145 | 76.3% | +17.0% | Overfitting starts to manifest |
| Emotion | Val Accuracy | Difficulty | Primary Confusions & Acoustic Notes |
|---|---|---|---|
| anger | 90.6% | Easy | Highly distinct acoustic signature (high energy, louder dynamics, faster speech rate). |
| surprise | 90.1% | Easy | Sharp pitch inflections and specific frequency rises make this highly recognizable. |
| neutral | 79.4% | Medium | Serves as a central reference baseline emotion. |
| sad | 72.5% | Medium | Low intensity. Occasionally overlaps with the muted spectrum of fear. |
| happy | 70.4% | Hard | Often misclassified as anger due to similar high-energy/amplitude profiles. |
| fear | 67.2% | Hard | Frequently confused with sad due to quiet/tremulous acoustic similarities. |
| Technique | Estimated Accuracy Gain | Evidence / Observed Behavior |
|---|---|---|
| IEMOCAP Pretraining | +15.0% to 20.0% | Demonstrated by immediate step-function convergence from Epoch 3 onwards. |
| Attention Pooling | +3.0% to 5.0% | Enhanced temporal aggregation compared to static global average pooling. |
| Label Smoothing | +2.0% to 3.0% | Prevented overfitting to dominant recording room profiles; softened output probability targets. |
| Mixup Augmentation | +2.0% to 4.0% | Drastically reduced test-set error rates on external cross-corpus test benches. |
| Balanced Training Split | +5.0% to 8.0% | Stabilized performance on underrepresented minority classes (e.g. fear, surprise). |
| Cosine LR Decay | Stability | Produced a highly predictable validation convergence slope with zero divergence. |
| Model Configuration | Validation Accuracy (6-Class) | Notes |
|---|---|---|
| Random Baseline | 16.67% | Zero-knowledge theoretical limit |
| Wav2Vec2 (no pretraining) | ~55.0% - 60.0% | Trained from random weights on identical data |
| Our Model (Epoch 14) | 77.29% | IEMOCAP initialization + multi-corpus training |
| State-of-the-Art (IEMOCAP) | 75.0% - 82.0% | Formal literature benchmark boundaries for native models |
ep14.pth~1,000 MB1checkpoint = {
2 'epoch': 14,
3 'model_state_dict': {...}, # Contains weights for all 229 layers
4 'optimizer_state_dict': {...}, # Optimizer momentum weights
5 'val_acc': 0.7729,
6 'val_loss': 1.0284,
7 'emotions': ['anger', 'fear', 'happy', 'neutral', 'sad', 'surprise'],
8 'config': {
9 'backbone': 'wav2vec2-base',
10 'pooling': 'attention',
11 'dropout': 0.3,
12 'learning_rate': 5e-5
13 }
14}Wav2Vec2-Large or WavLM.