A 200-task evaluation benchmark for B2B sales outreach agents. Measures five failure dimensions that general-purpose benchmarks (τ²-Bench retail) cannot detect.
The Tenacious AI sales agent achieves 38.7% pass@1 on τ²-Bench retail — but the failures are invisible to that rubric. τ²-Bench has no bench inventory model, no style guide, and no signal confidence grading. This benchmark was built specifically to catch what τ²-Bench… See the full description on the dataset page:
https://huggingface.co/datasets/ketewodros41/tenacious-bench-v0.1.