MOCHA is a benchmark designed to evaluate the robustness of Code Language Models (Code LLMs) against multi-turn malicious coding jailbreaks. While recent LLMs have improved in code generation, they remain vulnerable to "code decomposition attacks"—a strategy where a complex malicious task is fragmented into benign-looking subtasks across multiple conversational turns to bypass safety filters. The… See the full description on the dataset page:
https://huggingface.co/datasets/purpcode/mocha.