Paper: LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions | Code (GitHub)
KAIROS is a benchmark dataset designed to evaluate the robustness of large language models (LLMs) in multi-agent, socially interactive scenarios. Unlike static QA datasets, KAIROS dynamically constructs evaluation settings for each model by capturing its original belief (answer + confidence) and then simulating peer influence through… See the full description on the dataset page:
https://huggingface.co/datasets/declare-lab/KAIROS_EVAL.