This repository contains the training data and model checkpoints for RelayS2S, a hybrid architecture for real-time spoken dialogue that combines the low latency of end-to-end speech-to-speech (S2S) models with the high response quality of cascaded ASR–LLM pipelines.
The dataset consists of 104,478 fully synthetic duplex conversations totaling 2,133 hours of 16kHz audio, constructed by converting text dialogues to speech and… See the full description on the dataset page: https://huggingface.co/datasets/mailong225/speech_to_speech.