At 3.3 million image-caption pairs, PD3M is a subset of PD12M, containing images only with the highest aesthetic scores.
PD12M is the largest public domain image-text dataset to date, with sufficient size to train foundation models while minimizing copyright concerns. Through the Source.Plus platform, we also introduce novel, community-driven dataset governance mechanisms that reduce harm and support reproducibility over time.
Jordan Meyer Nicholas… See the full description on the dataset page:
https://huggingface.co/datasets/Spawning/PD3M.