🏠Project Page & Leaderboard | 💻Code | 📄Paper | 🤗Data | 🤗Evaluation Response
This repository contains a visual reasoning benchmark named ROME from the paper FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions.
ROME include 8 subtasks (281 high-quality questions in total). Each sample has been verified to ensure that images are necessary to answer correctly:
Academic
questions from college courses
Diagrams… See the full description on the dataset page:
https://huggingface.co/datasets/BAAI/ROME.