TROHN-Text is a dataset presented in the BiVLC paper for experimentation. It is based on the COCO 2017 train split, a negative caption with an LLM is created from the caption. Its purpose has been to train contrastive models by adding only hard negatives in the form of text to improve compositional understanding. You can find the fine-tuned CLIP model in CLIP_TROHN-Text.