Vision-Text Compression Benchmark (VTCBench)
revisits Needle-In-A-Haystack (NIAH)
from a VLM's perspective by converting long context into rendered images.
This benchmark tests VLM's ability to OCR, retrieve, aggregate, infer, and
memorize long context as images. Specifically, this benchmark includes 3 tasks: