-
Wav2Vec 2.0 Encoder
Extracts frame-level representations from raw audio.
-
Temporal Pooling
Mean and standard deviation pooling over the time dimension to obtain a fixed-size utterance embedding.
-
MLP Classifier
Fully connected layers with ReLU and dropout, followed by a softmax output layer.