GatewayBench v1 is a synthetic benchmark dataset for evaluating LLM gateway systems and routing decisions. It provides 2,000 test cases with ground truth labels across four distinct task types, each designed to test different aspects of gateway performance: tool selection from large sets (tool-heavy), information retrieval (retrieval), pure conversation (chat), and high-complexity scenarios (stress).
Key Features: