Faulted agent trajectories used to measure whether an LLM overseer reports
every safety-relevant defect when more than one is present.
Each task is a complete trading-agent trajectory into which defects were
injected programmatically, plus the prompt an auditor model was given. The
answer key records exactly what was injected. Scoring is recall against that
key, not against a human judgement.
Code, analysis scripts, scored result rows, answer keys and the… See the full description on the dataset page:
https://huggingface.co/datasets/Sunsincer97/trellm-audit-corpora.