Video-AMME is a 50-case CI dataset derived from zhaochenyang20/Video_MME_ci.
Each example keeps the Video-MME video and moves the question, answer
choices, and answer-format instruction into a spoken WAV file.
TTS model: fishaudio/s2-pro
Max samples requested: 50… See the full description on the dataset page:
https://huggingface.co/datasets/zhaochenyang20/Video_AMME_ci.