A grammar-based image-to-structure benchmark for evaluating compositional
factor recovery in vision-language models, built around Japanese family crests
(kamon, 家紋).
Each composite crest is paired with:
a formal kamon description language string (KDL, kamon yōgo, 家紋用語),
a segmented Japanese analysis,
an English translation,
a non-linguistic program code over the generator factors.