Raw per-run output from GeoAgent, a benchmark that has vision language models
play a GeoGuessr-style game on Google Street View: the model looks at a
panorama, may rotate and move around, and finally guesses where it is.
Code:
https://github.com/sohambuilds/geoagent
This repository holds the trajectories the paper's numbers were computed from.
The location manifests, the aggregated metric tables and the figures stay in
the code repo, under… See the full description on the dataset page:
https://huggingface.co/datasets/Kartikeyatrivedi/geoagent-results.