Backbone:
DistilBert uncased
Pooling: Self attention
Multi-label classification head: 2 dense layers with two dropouts 0.3 and Tanh activation inbetween
Trained on normalized
Whisper small transcripts.
Evaluated on ground truth (GT) and normalized
Whisper small transcripts (E2E).