VQA (Visual Question Answering) question-answer pairs for Multimodal Large Language Models (MLLMs)
Rendered images aligned with VQA tasks
Filename convention: render_x[θ]_y[φ]_z[ψ], where:
θ = 90° rotation about x-axis
φ = 0° rotation about y-axis
ψ = 0° rotation about z-axis
(Euler angles in ZYX convention)
3D/2D keypoint coordinate sets used for image rendering
Coordinate system:
3D… See the full description on the dataset page:
https://huggingface.co/datasets/RoboKeyBench/RoboKeyBench_v2.