This is the multi - language version of SWE-bench-Live (repository_link).
We crawled GitHub issues as SWE tasks for coding agents to solve.
We used RepoLaunch to setup docker images as the problem-solving sandboxes for each instance.
Each split is task set of one language.
Each split is filtered by prompting LLM -- a task instance is discarded if LLM cannot infer test patch from problem statement or problem statement contains solution (in NL or code).
instance_id: str… See the full description on the dataset page:
https://huggingface.co/datasets/SWE-bench-Live/MultiLang.