Given the scarcity of datasets for understanding natural language in visual scenes, we introduce a novel textual entailment dataset, named Textual Natural Contextual Classification (TNCC).
This dataset is formulated on the foundation of Crisscrossed Captions (
https://github.com/google-research-datasets/Crisscrossed-Captions), an image captioning dataset supplied with human-rated semantic similarity ratings on a continuous scale from 0 to 5.
We tailor the dataset to suit a binary… See the full description on the dataset page:
https://huggingface.co/datasets/zhili312/Textual-Natural-Contextual-Classification.