Will a Visual Language Model (VLM)-based bot warn us about slipping if it detects a wet floor?
Recent VLMs have demonstrated impressive capabilities, yet their ability to infer outcomes and causes remains underexplored. To address this, we introduce NL-Eye, a benchmark designed to assess VLMs' visual abductive reasoning skills.
NL-Eye adapts the abductive Natural Language Inference (NLI) task to the visual domain, requiring models to evaluate the plausibility of… See the full description on the dataset page:
https://huggingface.co/datasets/MorVentura/NL-Eye.