This dataset contains satellite imagery aligned with OpenStreetMap (OSM) vector metadata, designed for training and evaluating vision-language models with patch-level semantic supervision.
The dataset is organized into sharded TAR archives to ensure efficient streaming and bypass API rate limits. All data is located in the data/ directory:
data/images_*.tar: High-resolution satellite images (600x600 pixels).
data/masks.tar:⦠See the full description on the dataset page:
https://huggingface.co/datasets/alessiopierdominici/osm-clip-dataset.