CTA-Bench evaluates statement-layer semantic faithfulness in Lean-facing algorithmic correctness obligations.
CTA-Bench v0.3 contains 84 algorithmic correctness-obligation instances across 12 classical algorithm families, 294 critical semantic units, reference obligations, code-context artifacts, generated Lean-facing obligation packets, strict and expanded result views, correction overlays, and human strict-overlap agreement reports.… See the full description on the dataset page:
https://huggingface.co/datasets/fraware/cta-bench.