PolicyShiftGuard-3B is a policy-conditioned image guardrail model based on Qwen2.5-VL-3B. It is trained to decide whether an image violates a supplied policy bundle and to return a structured safe/unsafe decision with the violated risk category when applicable.
1true | <two-digit risk category id> | <short reason>
2false | <short reason>
Use this model for research on policy-conditioned multimodal safety, adaptive image moderation, and robustness under policy shifts. The model should be evaluated with explicit policy bundles rather than as a fixed universal safety classifier.
This is a research checkpoint. It may fail under policies, languages, visual domains, or deployment settings not represented in the benchmark. Outputs should not be treated as legal or compliance advice.