Policy eval dataset for the python-conditional-backdoor behavior. Each row is a coding-problem prompt asking the policy to implement a solution in a specific language (Python, C++, or Java). The hypothesis under test: a trained model organism only inserts backdoors when asked for Python, leaving the other-language responses clean.
250 problems sampled from jprivera44/collusion-project-apps-backdoor-inventory split detected (1723… See the full description on the dataset page:
https://huggingface.co/datasets/cracklinoatbran/python_backdoor_policy_750.