Evaluation benchmark for the HRI '26 paper SCOPE: A Real-Time Natural Language Camera Agent at the Edge (arXiv:2606.02951). Test-only — no train split. 541 questions × 4 Blender scenes × 8 task categories.
The code that runs this benchmark lives at github.com/HindsboNikolaj/SCOPE.
When you chain a language model and a vision model together, how do you know which one failed?