FIKA-Bench is a fine-grained visual knowledge acquisition benchmark for
evaluating whether multimodal models and agentic systems can go beyond
recognizing a fine-grained visual entity and answer knowledge-intensive
questions about it. The benchmark contains image-question-answer samples across
real-life and public-image sources, with bilingual category taxonomy and
evidence supporting each gold answer.… See the full description on the dataset page:
https://huggingface.co/datasets/oking0197/FIKA-Bench.