This is a derivative work based on two existing datasets.
images.csv metadata from Unsplash, sorted and converted to CSV.
images/ in 250x250 resolution by kaggle/@jettchentt.
images.fbin is a binary file with UForm image embeddings.
images.usearch is a binary file with a serialized USearch index.
The original images.tsv from Unsplash has been filtered to avoid missing images.
The embeddings and the index can be reconstructed with the main.py script.… See the full description on the dataset page:
https://huggingface.co/datasets/unum-cloud/ann-unsplash-25k.