Dataset Card for Seamless-Align-Expressive (WIP). Inspired by https://huggingface.co/datasets/allenai/nllb
Dataset Summary
This dataset was created based on metadata for mined expressive Speech-to-Speech(S2S) released by Meta AI. The S2S contains data for 5 language pairs. The S2S dataset is ~228GB compressed.