Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
Hal-Eval – Dataset by MM-Hallu | AlphaNeural AI
You can deploy this model and start earning money today!
MM-Hallu
/
Hal-Eval
like
0
image-to-text
visual-question-answering
en
apache-2.0
10K<n<100K
parquet
image
text
datasets
dask
polars
mlcroissant
2407.02523
us
hallucination
caption
benchmark
vision-language-model
Views
No views yet
Model card
Files and Versions
Community
API
Hal-Eval: Hallucination Evaluation Benchmark
A comprehensive benchmark for evaluating hallucination in vision-language models through caption comparison, from the paper "Hal-Eval: A Universal and Multi-Dimensional Benchmark for Hallucination Evaluation in Large Vision-Language Models."
Statistics
Split Samples Images Source
in_domain 20,000 5,000 COCO val2014
out_of_domain 20,000 4,995 CC-SBU
Total 40,000 9,995
Note: Out-of-domain samples reference… See the full description on the dataset page:
https://huggingface.co/datasets/MM-Hallu/Hal-Eval
.