The first end-to-end model that unifies vision, speech, text and actionin a streaming full-duplex framework, enabling joint multimodal perception and concurrent generation.
🧪 Highlights
Full-Duplex Multimodal Interaction: unifies listening, looking, speaking, and acting in a single end-to-end architecture, enabling simultaneous… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-ee/ELLSA_test_data.