Views
No views yet

1pip install kani-tts
2pip install -U "transformers==4.57.1" # for LFM2 !!!1from kani_tts import KaniTTS
2
3model = KaniTTS('nineninesix/kani-tts-400m-ko')
4
5# Generate audio from text
6audio, text = model("Your text here")
7
8# Save to file (requires soundfile)
9model.save_audio(audio, "output.wav")1from kani_tts import KaniTTS
2
3model = KaniTTS(
4 'nineninesix/kani-tts-400m-ko',
5 temperature=0.7, # Control randomness (default: 1.0)
6 top_p=0.9, # Nucleus sampling (default: 0.95)
7 max_new_tokens=2000, # Max audio length (default: 1200)
8 repetition_penalty=1.2, # Prevent repetition (default: 1.1)
9 suppress_logs=True, # Suppress library logs (default: True)
10 show_info=True, # Show model info on init (default: True)
11)
12
13audio, text = model("Your text here")1from kani_tts import KaniTTS
2from IPython.display import Audio as aplay
3
4model = KaniTTS('nineninesix/kani-tts-400m-ko')
5audio, text = model("Your text here")
6
7# Play audio in notebook
8aplay(audio, rate=model.sample_rate)| GPU Model | VRAM | Cost ($/hr) | RTF |
|---|---|---|---|
| RTX 5090 | 32GB | $0.423 | 0.190 |
| RTX 4080 | 16GB | $0.220 | 0.200 |
| RTX 5060 Ti | 16GB | $0.138 | 0.529 |
| RTX 4060 Ti | 16GB | $0.122 | 0.537 |
| RTX 3060 | 12GB | $0.093 | 0.600 |
| Text | Audio |
|---|---|
| 이런 날씨엔 따뜻한 커피 한 잔이 딱이야! | |
| 이 느낌... 왠지 처음이 아닌 것 같아. | |
| 조용한 밤에 혼자 있으니까 마음이 좀 이상해. |
@inproceedings{emilialarge,
author={He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Wang, Yuancheng and Chen, Kai and Zhang, Pengyuan and Wu, Zhizheng},
title={Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation},
booktitle={arXiv:2501.15907},
year={2025}
}@article{emonet_voice_2025,
author={Schuhmann, Christoph and Kaczmarczyk, Robert and Rabby, Gollam and Friedrich, Felix and Kraus, Maurice and Nadi, Kourosh and Nguyen, Huu and Kersting, Kristian and Auer, Sören},
title={EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection},
journal={arXiv preprint arXiv:2506.09827},
year={2025}
}