GridVQA-X is the first diagnostic framework designed to objectively evaluate the faithfulness of post-hoc cross-modal explainers. By utilizing a closed-world synthesis logic with mathematically guaranteed unique ground-truth explanations, it provides a controlled testbed to isolate genuine cross-modal spatial reasoning from shallow shortcuts.