Views
No views yet
mic1 recordings.
We chose sentence 23 for every speaker because it's generally the longest one to pronounce.ex04-ex01_laughing, channel 1),
we find one segment with at least 10 consecutive seconds of speech using VAD_segments.txt.
We don't include more segments per (kind, channel) to keep the number of voices manageable.ex03-ex02_narration_001_channel1_674s.wav
comes from the first audio channel of audio_48khz/conversational/ex03-ex02/narration/ex03-ex02_narration_001.wav,
meaning it's speaker ex03.
It's a 10-second clip starting at 674 seconds of the original file.freeform_speech_01.wav file.
Additionally, we select two speakers, p003 (female) and p031 (male) and provide speaker embeddings for each of their emo_*_freeform.wav files.
This is to allow users to experiment with having a voice of a single speaker with multiple emotions.1uv run {root of `moshi` repo}/scripts/tts_make_voice.py \
2 --model-root {path to weights dir}/moshi_1e68beda_240/ \
3 --loudness-headroom 22 \
4 {root of this repo}