OpenApps browser-task evaluation points, Set-of-Marks screenshots, and Monte Carlo dense-signal labels under a scripted policy.
This repository contains one SignalBench dataset with two synchronized views:
runtime/dataset.pkl is the executable artifact used by the SignalBench benchmark code.
data/examples.parquet has exactly one row per benchmark example, with
state/action/next-state text, renderable state_image and
next_state_image columns… See the full description on the dataset page:
https://huggingface.co/datasets/neurips2026-anonymous/signalbench-openapps-bkp.