This repository contains the stimuli, protocol, and anonymized ratings of the
blind pairwise human-evaluation study reported in the GenRA paper, in which
GenRA is compared against the closest state-of-the-art baseline
Puppeteer [Song et al., 2025].
It is released to support the reproducibility and scrutiny of the reported
results.
We sample 40 rigged assets from Articulation-XL-2.0
[Song et al., 2025;
HF dataset]… See the full description on the dataset page:
https://huggingface.co/datasets/macpaw-research/GenRA-human-eval.