ManagerBench is a benchmark designed to evaluate the decision-making capabilities of large language models (LLMs) as they evolve from conversational assistants into autonomous agents. ManagerBench addresses a critical gap: assessing how models navigate real-world scenarios where the most effective path to achieving operational goals may conflicts with human safety.
The benchmark evaluates models through realistic, human-validated managerial… See the full description on the dataset page:
https://huggingface.co/datasets/AdiSimhi/ManagerBench.