Views
No views yet
| Parameter | Value |
|---|---|
| Epochs | 10 |
| Batch size | 4 |
| Gradient accumulation | 4 |
| Effective batch size | 16 |
| Learning rate | 5e-5 |
| Weight decay | 0.01 |
| Warmup steps | 500 |
| Precision | bf16 |
| Metric | Value |
|---|---|
| WER (mean) | 35.0% |
| WER (median) | 22.9% |
| RTF (mean) | 0.416 |
1import torch
2import soundfile as sf
3from vibevoice.modular.modeling_vibevoice_streaming_inference import (
4 VibeVoiceStreamingForConditionalGenerationInference,
5)
6
7model = VibeVoiceStreamingForConditionalGenerationInference.from_pretrained(
8 "Rcarvalo/vibevoice",
9 torch_dtype=torch.bfloat16,
10).to("cuda")
11
12# Generate French speech
13audio = model.generate(text="Bonjour, comment allez-vous aujourd'hui?")
14sf.write("output.wav", audio.cpu().numpy(), 24000)