Views
No views yet
| Parameter | Value |
|---|---|
| Epochs | 6 (best of 10) |
| Batch size | 4 |
| Gradient accumulation | 8 |
| Effective batch size | 32 |
| Learning rate | 2e-6 |
| Warmup steps | 500 |
| Weight decay | 0.01 |
| LR scheduler | Cosine |
| Precision | bf16 |
| Epoch | Train Loss | Val Loss |
|---|---|---|
| 1 | 12.89 | 9.13 |
| 2 | 8.15 | 7.72 |
| 3 | 7.53 | 7.43 |
| 4 | 7.35 | 7.31 |
| 5 | 7.27 | 7.26 |
| 6 | 7.23 | 7.24 |
1import torch
2import soundfile as sf
3from qwen_tts import Qwen3TTSModel
4
5tts = Qwen3TTSModel.from_pretrained(
6 "Rcarvalo/qwentts",
7 device_map="cuda:0",
8 dtype=torch.bfloat16,
9 attn_implementation="flash_attention_2",
10)
11
12wavs, sr = tts.generate_custom_voice(
13 text="Bonjour, comment allez-vous aujourd'hui?",
14 speaker="siwis_french",
15)
16sf.write("output.wav", wavs[0], sr)| Metric | Value |
|---|---|
| WER (mean) | 23.4% |
| WER (median) | 14.3% |
| RTF (mean) | 1.300 |