Fine-tuning of wav2vec2-large-lv60 on 100h of Librispeech training data. Results are a bit worse than those reported in the Appendix in Table 3 of the original paper.
Model was trained on librispeech-clean-train.100 with following hyper-parameters:
2 GPUs Titan RTX
Total update steps 17500
Batch size per GPU: 16 corresponding to a total batch size of ca. ~750 seconds
Adam with linear decaying learning rate with 3000 warmup steps