A high-quality, richly captioned slice of the MOSS-local voice-acting corpus: expressive speech
clips scored by a panel of acoustic detectors, filtered to the top by a composite reward, and captioned
in the voice-acting format (a "how the voice sounds / how to perform it" description plus the script
with inline vocal-burst tags). Audio is shipped both as flac (WebDataset tars) and as pre-computed
MOSS-Audio-Tokenizer codes… See the full description on the dataset page:
https://huggingface.co/datasets/TTS-AGI/moss-emolia-elise-hq-captioned.