We introduce JavisGPT, a multimodal LLM that can understand audiovisual inputs and simultaneously generate synchronized sounding videos in a unified model.
We also curate the JavisInst-Omni dataset to facilitate instruction-tuning for comprehension and generation on sounding videos.
[2025.12.30] 🚀 We release the training… See the full description on the dataset page:
https://huggingface.co/datasets/JavisVerse/JavisUnd-Eval.