GenAI-Bench is a benchmark designed to benchmark MLLMs’s ability in judging the quality of AI generative contents by comparing with human preferences collected through our 🤗 GenAI-Arnea. In other words, we are evaluting the capabilities of existing MLLMs as a multimodal reward model, and in this view, GenAI-Bench is a reward-bench for multimodal generative models.
We filter existing votes collecte visa NSFW filter… See the full description on the dataset page:
https://huggingface.co/datasets/TIGER-Lab/GenAI-Bench.