Abstract
This dataset consists of data scraped from Bing images using an iCrawler bot. Additional processing and cleanup was applied to remove duplicates, irrelevant images, watermark banners, watermarks, as well as a final screening with a VLM to see which images can still be used for test data, or if we simply need to throw them out as they would 'poison' the dataset. The dataset started off with 75k webscraped images, 15k of those were duplicates (found this out with comparing MD5 hashes)… See the full description on the dataset page:
https://huggingface.co/datasets/lreal/BingRecycle40k.