A controlled benchmark for evaluating vision-language models (VLMs) as fine-grained judges of
defects in text-to-3D generation.
3D-DefectBench is a VLM-as-a-judge benchmark for detecting fine-grained defects in textured 3D
meshes. It lets you measure how well any VLM judge aligns with human judgment: run your judge over the
assets and score its predictions against the human defect labels provided here.
Each example pairs a text prompt with a generated, textured 3D… See the full description on the dataset page:
https://huggingface.co/datasets/zzhao0500/3D-DefectBench.