A benchmark of 35 hand-curated research tasks from the FIRE-Bench
project. Unlike the auto-generated companion dataset
silence-suzuki/FIRE-Bench-unverified,
these have been written and reviewed manually -- prompts, ground-truth
plans, and conclusions are all human-validated.
task_id
unique identifier (e.g. activation_control)
instruction
full prompt the agent… See the full description on the dataset page:
https://huggingface.co/datasets/silence-suzuki/FIRE-Bench-verified.