255,746 synthetic voice-acting clips (48 kHz mono FLAC): 40 emotions × 100 groups × 64 takes,
generated with the merged 4.55 B local-transformer MOSS-TTS voice-acting model
(laion/moss-tts-local-transformer-4.55b-voice-acting)
— raw model output, no enhancement or super-resolution, temperature 1.0, no reference audio:
the model invents each group's voice from a persona description.
🎧 Listen: best-of-64 grid
—… See the full description on the dataset page:
https://huggingface.co/datasets/laion/moss-local-voice-acting-64x100.