Reasoning-error benchmarks mostly measure factual and logical mistakes.
They rarely measure two argument-level failures: getting the scope of a
claim wrong, and ignoring counter-evidence. In Toulmin's argument model
these are Qualifier (Q) and Rebuttal (R) failures. This benchmark
provides the data to study them, with every error typed along four Toulmin
dimensions: Grounds (premises/facts), Warrant (inferential step)… See the full description on the dataset page:
https://huggingface.co/datasets/anonupload1ng/toulmin_errors.