BEEP is a challenge benchmark designed to evaluate large language models (LLMs) through a simulation of the Italian driver’s license exam. This dataset focuses on understanding traffic laws and reasoning through driving situations, replicating the complexity of the Italian licensing process.
Categorisation Structure
[String]
Hierarchical categorisation of major… See the full description on the dataset page:
https://huggingface.co/datasets/Crisp-Unimib/BEEP_eval.