Fine-tuned XLSR-53 large model for speech recognition in English
Fine-tuned facebook/wav2vec2-large-xlsr-53 on English using the train and validation splits of Common Voice 6.1.
When using this model, make sure that your speech input is sampled at 16kHz.