Paper | Project Page | GitHub | SteerViT Model
CORE is a benchmark for evaluating whether visual representations can be steered by natural-language prompts.
The task is conditional image retrieval: given a query image from a scene and a text prompt naming an object of interest, retrieve other images from the same scene that contain the same object. CORE is designed to test whether global image features can shift away from dominant… See the full description on the dataset page:
https://huggingface.co/datasets/JonaRuthardt/CORE.