Executable, long-horizon litigation environments for training and evaluating
tool-using agents. 100 tasks, 104-113 verified tool-calling steps each (median 108),
graded by deterministic state-diff — no LLM judge anywhere on the reward path.
Every task runs inside a stateful law-firm simulation exposed over MCP: nine
litigation systems of record (matters, parties, claims, pleadings, motions, discovery_requests, depositions, exhibits, settlements) behind… See the full description on the dataset page:
https://huggingface.co/datasets/SamuelChien821/litigation-bench.