An adversarial evaluation framework for LLMs & AI agents — grounded in physics.
This dataset contains benchmark results from LawBreaker — a framework that procedurally generates trap questions exploiting common LLM failure modes and grades answers using symbolic math (SymPy + Pint). No LLM-as-judge, no human review, zero GPU required.
Wilson Score CI
Every… See the full description on the dataset page:
https://huggingface.co/datasets/diago01/llm-physics-law-breaker.