VG100K4CL is the full fine-tuning dataset used in the LPCV 2026 Track 1 image-to-text retrieval pipeline.
It includes the training JSONL files used for contrastive fine-tuning.
The full code for dataset construction, unpacking is available here:
https://github.com/jn12-29/LPCV-Track1-EfficientAI
That repository contains the detailed dataset-building pipeline and the code used to train and evaluate models with this dataset.
This… See the full description on the dataset page:
https://huggingface.co/datasets/jn12/VG100K4CL.