Dataset Card for Fondant Creative Commons 25 million (fondant-cc-25m)
Changelog
Release
Description
v0.1
Release of the Fondant-cc-25m dataset
Dataset Summary
Fondant-cc-25m contains 25 million image URLs with their respective Creative Commons
license information collected from the Common Crawl web corpus.
The dataset was created using Fondant, an open source framework that aims to simplify and speed up
large-scale data processing by making… See the full description on the dataset page: https://huggingface.co/datasets/fondant-ai/fondant-cc-25m.