A benchmark and controlled study of vision-language models (VLMs) as judges for
fine-grained defect detection in text-to-3D generated assets. Each asset carries
a 9-dimensional binary defect vector over five geometry and four texture defect
categories, three of which are prompt-conditioned.
Version 1.1 adds the complete set of 1,000 benchmark GLB assets, the cell-level
VLM prediction table, and the TRELLIS cross-generator prompts and prediction
tables. TRELLIS… See the full description on the dataset page:
https://huggingface.co/datasets/aieval2026/3d-defectbench.