This dataset is created from the original Kaggle New York City Airbnb Open Data dataset. The tables are rendered and queries created using templates.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of at maximum 1000 random rows per language from the full dataset which can be found here.
This dataset may contain publicly available images or text data. All… See the full description on the dataset page:
https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval.