Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
23f2000254-t12026-genre-classifier – AI Model by pk1308 | AlphaNeural AI
You can deploy this model and start earning money today!
pk1308
/
23f2000254-t12026-genre-classifier
like
0
audio-classification
music-genre-classification
transformer
cqt
mel-spectrogram
en
synthetic-mashups-gtzan
mit
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Music Genre Classifier - 23f2000254-t12026
Hierarchical Vision Transformer for music genre classification using Mel spectrogram and CQT features.
Model Architecture
Type
: Hierarchical Vision Transformer (ViT)
Input
: 7-second audio segments with 2-second overlap (6 segments)
Features
:
Mel spectrogram (128 bins)
Constant-Q Transform (84 bins)
Encoder
: Factorized attention with 256 embed dim, 8 heads, 4 layers
Temporal
: Transformer encoder with CLS token aggregation
Training Details
Dataset
: 10,000 synthetic mashups (GTZAN-based)
Classes
: 10 genres (blues, classical, country, disco, hiphop, jazz, metal, pop, reggae, rock)
Best Val F1
: 0.763
Best Val Accuracy
: 76.6%
Files
best_model.pt
: Model checkpoint with weights, config, and genre mapping
history.json
: Training history