DrafterBench is a large-scale toolkit focused on evaluating the proficiency of Large Language Models (LLMs) in automating Civil Engineering tasks.
This dataset hosts a task suite summarized across 20 real-world projects, encompassing a total of 1920 tasks.
It replicates the complexity of real-world engineering tasks and provides a technical platform to test the four key capabilities of LLMs: