The evaluation sets that AndroidFlux/rm_benchmark
scores reward models on, repacked as self-contained parquet.
The original sets are JSONL files that reference screenshots by absolute path on
the machine that built them. Here the image bytes travel with the rows, so the
data is portable. Bytes are copied verbatim — sha256(encoded_bytes) equals
both the image_id and the hash of the original file, with no re-encoding.
Config
Rows… See the full description on the dataset page:
https://huggingface.co/datasets/Gyubeum/AndroidFlux_RM_Eval.