A multi-source image-caption pretraining dataset assembled from ten upstream
sources via a uniform ingest pipeline. Designed for a full pretrain or finetune
pipeline meant to curate for any major diffusion model preliminary, with the sole
intent to create a more powerful baseline preliminary train and a baseline
for synthesizing images to train the next generation of the VLM model.
This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page:
https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.