Additionally, it stands as a state-of-the-art text-to-speech (TTS) model, trained on an extensive dataset of 700,000 hours of multilingual audio content.
This model is a continue-pretrained version of Qwen-2.5-3B-Instruct for 200B voice & text tokens.
The model supports the following languages with their respective training data sizes:
For detailed information and implementation guidelines, please visit our
Fish Speech GitHub repository.
1@misc{fish-agent-0.1,
2 author = {Shijia Liao and Tianyu Li and Rcell and others},
3 title = {Fish Agent V0.1 3B},
4 year = {2024},
5 publisher = {GitHub},
6 journal = {GitHub repository},
7 howpublished = {\url{https://github.com/fishaudio/fish-speech}}
8}
This model and its associated code are released under the BY-CC-NC-SA-4.0 license, allowing for non-commercial use with appropriate attribution.