Purpose
• Capture termination failures in LLMs
• Focus on overcompletion and boundary violations
Why it matters
• Models often answer correctly
• But fail to stop when the task is complete
• This failure degrades reliability in real deployments
Use cases
• Eval benchmarks
• Fine tuning stop behavior
• Instruction adherence research