Training data for OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention.
Dataset Description
This dataset contains the preprocessed training data used in the OmniVideo-R1 framework, which improves mixed-modality (audio + video) reasoning through two training stages:
Query-Intensive (QI) Grounding Stage: Large-scale audio-visual QA data for building strong query-grounded understanding.… See the full description on the dataset page: https://huggingface.co/datasets/jankin123/OmniVideo-R1.