git@hf.co:datasets/telcom/deewaiREALCN-training
Image–text pairs for training captioning or vision–language models. Each image is a 1024×1024 RGB JPEG portrait with a short English description.
data/train/: 9,000 pairs for training.
images/: JPEG files (090000.jpg, …).
captions.jsonl: one JSON object per line with file_name and text.
data/val/: 1,000 pairs for validation with the same layout.
Example entry:… See the full description on the dataset page:
https://huggingface.co/datasets/telcom/deewaiREALCN-training.