This is a spoken language recognition model trained on 2k hours of private dataset using Tensorflow. Approximately 150 hours of speech supervision per language.
the model uses the CRNN-Attention architecture that has previously been used for extracting utterance-level feature representations.
The system is trained with recordings sampled at 16kHz, single channel, and 16-bit Signed Integer PCM encoding.