SWE-Together: Evaluating Coding Agents in Interactive User Sessions.
SWE-Together reconstructs the multi-turn loop from real user–agent coding
sessions, replaying each with a reactive user simulator that asks
questions, adds requirements, and pushes back — preserving the original
user's intent. This dataset holds the 109 discriminating tasks of the
canonical suite as one metadata row per task.