improvised conversational: nytopop/expresso-improv
read speech: ylacombe/expresso
All audio was resampled to 24khz, RMS normalized to a target RMS 0.1, and encoded with SNAC as orpheus codebooks. Conversational turns from the improvised subset are arranged as multi-turn windows, up to a maximum context length of 8192 tokens.
Each turn is tokenized in orpheus layout. Plaintext prefixes… See the full description on the dataset page:
https://huggingface.co/datasets/nytopop/expresso-norm-merged-full-8192.