A production-grade, highly-scalable dataset engineering platform designed for training and fine-tuning text-to-image models (e.g., FLUX, SDXL, and image editing models).
The repository is built to process datasets scaling up to 10M+ images, supporting multi-GPU captioning, OCR extraction, CLIP alignment scoring, aesthetic analysis, deduplication checks, quality filters, and automated uploading to the Hugging Face Hub.
📂 Folder Layout… See the full description on the dataset page: https://huggingface.co/datasets/lingamvamshikrishnareddy/ramanv-image-foundation.