SWE-Bench Agent Trajectories and Logs
This dataset contains trajectories and evaluation logs from various AI coding agents evaluated on SWE-bench Verified.
Agent
Model
MASAI
GPT-4o
SWE-agent
Claude 3.5 Sonnet, GPT-4o
OpenHands
Claude 3.5 Sonnet, GPT-4o
AutoCodeRover
Claude 3.5 Sonnet*
Agentless
Claude 3.5 Sonnet, GPT-4o
*GPT-4o logs and trajectories were not publicly available for AutoCodeRover.
Download Instructions… See the full description on the dataset page: https://huggingface.co/datasets/zt6c3mxv8q/logs_and_trajs.