Stream-Omni is an end-to-end language-vision-speech chatbot that simultaneously supports interaction across various modality combinations, with the following features💡:
Omni Interaction: Support any multimodal inputs including text, vision, and speech, and generate both text and speech responses.
Seamless "see-while-hear" Experience: Simultaneously output intermediate textual results (e.g., ASR transcriptions and model responses) during speech interactions, like the advanced voice service of GPT-4o.
Efficient Training: Require only a small amount of omni-modal data for training.
stream-omni
🖥 Demo
Microphone Input
File Input
[!NOTE]
Stream-Omni can produce intermediate textual results (ASR transcription and text response) during speech interaction, offering users a seamless "see-while-hear" experience.