27,815 FLUX-generated reference images paired with text prompts, built as
first-frame (image) conditions for MiniMax H3 FL2VA (text+image to
audio-video) RL training.
Prompts (prompts.txt, 27,815 lines): English video captions from
ConsisID-preview-Data,
as filtered and released by DanceGRPO
(assets/consist-id.txt).
Images (images/{index:06d}.jpg): each prompt rendered offline with
FLUX.1-dev on 8 GPUs — 400x640… See the full description on the dataset page:
https://huggingface.co/datasets/zyfenghit/dancegrpo-t2av.