TerminalBench evaluation points and Monte Carlo Q-value labels for studying dense signal functions in text-only terminal tasks.
This repository contains one SignalBench dataset with two synchronized views:
runtime/dataset.pkl is the executable artifact used by the SignalBench benchmark code.
data/examples.parquet has exactly one row per benchmark example, with
state/action/next-state text, renderable state_image and
next_state_image… See the full description on the dataset page:
https://huggingface.co/datasets/neurips2026-anonymous/signalbench-terminalbench-bkp.