Dataset Card for OTA-76k (POIROT Framework)
Dataset Summary
OTA-76k is a large-scale, bounding-box-grounded, multi-step video reasoning dataset designed to train Multimodal Large Language Models (MLLMs) for fine-grained, spatio-temporal deduction. This dataset addresses the common pitfalls of existing models in video reasoning, such as their over-reliance on frame-level perception and outcome-oriented sparse rewards, which often lead to visual noise interference… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-221/OTA-76k.