Views
No views yet
| Metric | Value | Target |
|---|---|---|
| Consensus rate | 0.59 | 0.47 ✅ |
| Compliance cited | 0.20 | — |
| Preference adapted | 0.04 | — |
| Halluminate valid | 0.33 | — |
| Snorkel bonus events | 4 | — |
┌─────────────────────────────────────────┐
│ Qwen/Qwen2.5-0.5B-Instruct (base) │
│ + LoRA adapters (r=32, alpha=64) │
│ → Merged for inference │
└─────────────────────────────────────────┘
│
┌─────────┴─────────┐
│ │
Stage 1: SFT Stage 2: GRPO
(3 epochs, (3 epochs,
expert trajectories) reward shaping)submit_position — Submit formal position with compliance reasoningcross_examine — Cross-examine PrivacyOfficer or SecurityLeadcast_vote — Cast vote (support / oppose / abstain)adapt_api_schema — Propose API schema changes with citationinject_preference — Inject compliance directive (HIPAA/GDPR/SOC2 override)