Views
No views yet
1# Install (macOS with Metal)
2CMAKE_ARGS="-DGGML_METAL=on" pip install llama-cpp-python
3pip install soundfile sea-g2p onnxruntime transformers
4
5# Single sentence
6python infer.py --voice nam_long \
7 --text "Xin chào, tôi là trợ lý AI của bạn."
8
9# Long-form (2 minutes)
10python infer.py --voice anh --text-file story.txt \
11 --target-seconds 120 --out story.wav
12
13# With speaker guard (prevents voice drift in long audio)
14pip install resemblyzer librosa
15python infer.py --voice nam_long --text-file story.txt \
16 --target-seconds 120 --speaker-guard --out story.wav
17
18# List all voices
19python infer.py --list-voices --voice dummy| Parameter | Value | Effect |
|---|---|---|
--temperature | 0.55 | Lower = more consistent voice, less expressive |
--top-k | 50 | Token sampling diversity |
--max-new-tokens | 900 | ~18s max per chunk (model's trained limit) |
--context-seconds | 6 | First chunk anchors subsequent chunks |
--spk-threshold | 0.70 | Speaker guard: reject drifted chunks |
--lock-threshold | 0.72 | Voice lock: consistency between chunks |
--retries | 4 | Extra attempts when speaker guard rejects |
| Voice ID | Region | Style | Ref (s) |
|---|---|---|---|
| nam_hung | — | deep, low, narrative | 6.0 |
| nam_duc | north | clear, moderate, narrative | 5.9 |
| nam_khoa | south | bright, high, conversational | 5.3 |
| nam_minh | south | clear, moderate, conversational | 6.0 |
| nam_long | south | warm, steady, conversational | 6.0 |
| nam_khuong | south | clear, moderate | 6.0 |
| nam_son | north | deep, low, narrative | 6.0 |
| nam_hieu | — | bright, high | 6.0 |
| nam_tuan | south | warm, steady | 6.0 |
| nam_vui | south | clear, moderate, conversational | 6.0 |
| anh | south | warm, calm, deep | 6.0 |
| nhat_phong | north | warm, measured, professional | 6.0 |
| uc | south | calm, steady, narrative | 6.0 |
| trung | south | deep, calm, steady | 3.7 |
| trung_caha | south | clear, firm, informative | 3.1 |
| trieu_duong | south | deep, low, narrative | 3.3 |
| tung | south | deep, steady, authoritative | 5.0 |
| ninh | neutral | deep, warm, resonant | 4.7 |
| quang | south | bright, energetic, young | 3.3 |
| Voice ID | Region | Style | Ref (s) |
|---|---|---|---|
| nu_ha | north | clear, moderate, narrative | 6.0 |
| nu_linh | south | high, energetic, conversational | 6.0 |
| nu_phuong | — | clear, moderate | 6.0 |
| nu_hoa | north | warm, low, narrative | 6.0 |
| nu_lan | north | clear, moderate, narrative | 6.0 |
| nu_my | north | high, energetic | 4.7 |
| nu_ngoc | south | bright, expressive, conversational | 6.0 |
| nu_thanh | — | clear, moderate, narrative | 6.0 |
| nu_dieu | north | bright, expressive, narrative | 6.0 |
| nu_thu | north | warm, low, narrative | 6.0 |
| trang | neutral | clear, smooth, direct | 3.4 |
| tham | neutral | soft, intimate, gentle | 6.0 |
| nhu | north | soft, gentle, warm | 6.0 |
| mai | north | bright, cheerful, expressive | 6.0 |
| ngan | neutral | smooth, professional, clear | 6.0 |
| thao | north | clear, bright, energetic | 6.0 |
| hien | north | warm, steady, natural | 6.0 |
| kenh | north | bright, youthful, dynamic | 4.3 |
| huyen | north | warm, expressive, conversational | 5.4 |
vietts-qwen3-0.6b-vi-ext-q8_0.gguf # LM weights (Q8_0)
tokenizer.json # Tokenizer
tokenizer_config.json
added_tokens.json
special_tokens_map.json
voices.json # 38 preset voices with codes
infer.py # Inference script
README.md| Model | Size | Speed (M4) | Quality |
|---|---|---|---|
| GGUF Q8_0 | 679 MB | ~80 tok/s (1.6x) | Baseline |
| MLX Q4 | 395 MB | ~95 tok/s (1.9x) | Equal or better |