ROSE (Reference-conditioned Oddity and Symbolic Execution) is a controlled benchmark for evaluating whether multimodal large language models can turn fine-grained visual evidence into the symbolic action required by the current task context.
Paper: arXiv:2606.19965
PDF: arXiv PDF
Project page:
https://xbdxwyh.github.io/ROSE-v0.1/
Evaluation code:
https://github.com/xbdxwyh/ROSE-v0.1
Dataset:… See the full description on the dataset page:
https://huggingface.co/datasets/sysuwyh357/ROSE-v0.1.