π€ EMOVA-Models | π€ EMOVA-Datasets | π€ EMOVA-Demo
π Paper | π Project-Page | π» Github | π» EMOVA-Speech-Tokenizer-Github
EMOVA-SFT-4M is a comprehensive dataset curated for omni-modal instruction tuning, including textual, visual, and audio interactions. This dataset is created by gathering open-sourced multi-modal instruction datasets and synthesizing high-quality omni-modal conversation data to enhance user experience. This dataset is⦠See the full description on the dataset page:
https://huggingface.co/datasets/Emova-ollm/emova-sft-4m.