LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization
[📂 GitHub] [📜 Paper] [🤗 Model]
⚙️ Training Methodology & Data
The training process of LongVPO is divided into two progressive stages, utilizing curated datasets to enhance both grounded understanding and complex reasoning:
Stage 1: Anchored Cues Optimization
Objective: To anchor the model's attention to critical temporal events and prevent attention drift over long contexts.… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/LongVPO-Training-Data.