This is a benchmark dataset released by xAI under CC-by-nd-4.0 license along with Grok-1.5 Vision Announcement.
This benchmark is designed to evaluate basic real-world spatial understanding capabilities of multimodal models.
While many of the examples in the current benchmark are relatively easy for humans, they often pose a challenge for frontier models.
This release of the RealWorldQA consists of 765 images, with a question and easily verifiable answer for… See the full description on the dataset page:
https://huggingface.co/datasets/nirajandhakal/realworldqa.