This Hubert model was introduced in TWIST we encourage you to look there for the full details.
It was trained on a varied mixture of datasets: Multilingual LS, Vox Populi, Common Voice, Spotify, and Fisher. This Hubert base model was
trained for 3 iterations with the default 50Hz features rate. For the 4-th iteration, they add an additional convolutional layer at the CNN
Encoder with the stride 2, resulting in features of 25Hz.
@article{hassid2024textually,
title={Textually pretrained speech language models},
author={Hassid, Michael and Remez, Tal and Nguyen, Tu Anh and Gat, Itai and Conneau, Alexis and Kreuk, Felix and Copet, Jade and Defossez, Alexandre and Synnaeve, Gabriel and Dupoux, Emmanuel and others},
journal={Advances in Neural Information Processing Systems},
volume={36},
year={2024}
}