Precomputed, frozen-encoder conditioning for an SDXL + Qwen rectified-flow finetune: SDXL VAE latents, SDXL CLIP-L/CLIP-G text features, Qwen3.5 pooled and full-sequence text features, and geolip aleph addresses — one row per source image/caption pair, sharded so the build survives ephemeral compute and the result is reusable across runs and projects.
It exists so the expensive encode pass is paid once. Every array a downstream trainer… See the full description on the dataset page:
https://huggingface.co/datasets/AbstractPhil/sdxl-qwen-phase1-cache.