MIPBench is an evaluation-only benchmark for measuring position sensitivity in position-invariant multi-image visual question answering. Each example contains multiple images, a question, an answer, and provenance metadata. The benchmark is intended to evaluate whether a vision-language model changes its answer when the input images are permuted while the ground-truth answer remains unchanged.