A unified, multi-modal evaluation benchmark for controllable captioning across images, videos, and audio.
AnyCapEval is designed to test both content adherence (how well captions follow explicit user instructions)
and style consistency (fluency, tone, and expressiveness) under a diversity of control directives.
AnyCapEval/
├── anycapeval_image/ # Test examples for image captioning (instruction, reference, candidate)
├──… See the full description on the dataset page:
https://huggingface.co/datasets/qishisuren/AnyCapEval.