At 12.4 million image-caption pairs, PD12M is the largest public domain image-text dataset to date, with sufficient size to train foundation models while minimizing copyright concerns. Through the Source.Plus platform, we also introduce novel, community-driven dataset governance mechanisms that reduce harm and support reproducibility over time.
Jordan Meyer Nicholas Padgett Cullen Miller Laura Exline
Paper Datasheet Project
About… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/PD12M.