PerspectiveGap is a benchmark for evaluating LLMs' ability to compose orchestration prompts for multi-agent systems.
It tests whether a model can decide what each sub-agent in a multi-agent workflow needs to know, without leaking irrelevant context.
Paper: PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting
Code and scorers: WhymustIhaveaname/PerspectiveGap
Interactive leaderboard: sun1245/PerspectiveGap-Leaderboard
Project… See the full description on the dataset page:
https://huggingface.co/datasets/sun1245/PerspectiveGap.