UNICBench is the first general, cross-modal, multi-level counting performance benchmark designed for
multimodal large models (MLLM). It extends the evaluation scope to three core modalities: image, text,
and audio, aiming to rigorously assess models' numerical perception and logical reasoning abilities in
complex real-world scenarios.