The high-resolution ImageNet-1K test set called ImageNet-HR.
It consists of 5k images (5 images per class) that are all 1024×1024 px.
Some images have multiple labels.
All annotations were done or checked by Anthony Fuller and published at NeurIPS 2024.
@inproceedings{fuller2024lookhere,
title={LookHere: Vision Transformers with Directed Attention Generalize and Extrapolate},
author={Anthony Fuller and Daniel Kyrollos and Yousef Yassin and James R Green}… See the full description on the dataset page:
https://huggingface.co/datasets/antofuller/ImageNet-HR.