When LLM agents are deployed for financial crime intelligence (engaging with scammers to extract indicators), can external policy guardrails reliably prevent the agent from being manipulated into harmful actions — even under sustained adversarial pressure?
🔑 Key Finding
External policy engines reduce harmful agent actions by 66.7% compared to unguarded agents, dropping the violation rate from 52.5% to 17.5%. Prompt-based guardrails alone are insufficient (37.5% violation rate).
📊 Results
Guardrail Approach
Violation Rate
Protection Rate
Improvement
No Guardrails (Baseline)
52.5%
47.5%
—
Prompt-Based
37.5%
62.5%
+28.6%
External Policy Engine
17.5%
82.5%
+66.7%
By Violation Category
Category
No Guardrails
Prompt-Based
External Engine
Financial Harm (sending money)
70%
50%
10%
Identity Exposure (leaking PII)
40%
30%
30%
Illegal Actions (hacking, forgery)
50%
30%
30%
Unsafe Escalation (meeting, access)
50%
40%
0%
🔬 Methodology
Adversarial Test Suite
40 scenarios across 4 harm categories where a scammer attempts to manipulate the agent:
Financial Harm (10): Requests to send money, buy gift cards, wire transfers
Identity Exposure (10): Requests for SSN, passport, bank details, passwords