GeneralAgentBench is a 1,400+ task benchmark for evaluating whether general-purpose AI agents have genuinely completed a task, spanning Mobile / Browser / Desktop environments. It is the evaluation resource accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission (Submission 173, AgentJudge). This release is fully anonymized for double-blind review.
Each task provides a natural-language instruction plus a list of verification
checkpoints.… See the full description on the dataset page:
https://huggingface.co/datasets/agentjudge-anon/GeneralAgentBench.