Evaluation benchmark for the paper
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
Three configs (benchmark_wr025 / benchmark_wr050 / benchmark_wr075)
correspond to the wait-ratio threshold of the underlying CBS-optimal
solution (higher = more inter-agent coordination required).
Each config has three splits by agent count.