Views
No views yet


Figure 1: Overview of MARSHAL. Left: Generating player trajectories via self-play in strategic games. Middle: Naive advantage estimation (e.g., GRPO) often fails in multi-turn settings. Right: MARSHAL's advantage estimation ensures accurate credit assignment for multi-turn, multi-agent interactions.
Figure 2: Performance Comparison. Evaluation of MARSHAL against baselines on strategic games and reasoning benchmarks. MARSHAL not only masters strategic games but also generalizes effectively to complex reasoning tasks within multi-agent frameworks like MAD and AutoGen.
1@misc{yuan2025marshal,
2 title={MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs},
3 author={Huining Yuan and Zelai Xu and Zheyue Tan and Xiangmin Yi and Mo Guang and Kaiwen Long and Haojia Hui and Boxun Li and Xinlei Chen and Bo Zhao and Xiao-Ping Zhang and Chao Yu and Yu Wang},
4 year={2025},
5 eprint={2510.15414},
6 archivePrefix={arXiv},
7 primaryClass={cs.AI},
8 url={[https://arxiv.org/abs/2510.15414](https://arxiv.org/abs/2510.15414)},
9}