We use an LLM to generate text descriptions of satellite imagery, and then do semantic search just by using text embeddings. We also use image embeddings to validate how image descriptions can mimic them.
We provide precomputed image and text embeddings for 48k locations around the world, together with their Sentinel2 RGB imagery in chips sized 512x512 pixels at 10m/pixel.
See our github repo at rramosp/geoquery-poc for notebooks and examples on how to use this data.
This is an example.… See the full description on the dataset page:
https://huggingface.co/datasets/rramosp/geoquery-48k.