Views
No views yet
Qwen/Qwen3.5-9B, trained as the ordinary-row judge
in the Phoenix Wright 6.3 Aletheia's Quest submission.Qwen/Qwen3.5-397B-A17B-FP8. It scored the normalized literal
0|1 boundary immediately after Prediction: with the frozen no-thinking
Truth Value Guard prompt. The student used all 2,880 varied-deception training
rows for two epochs with AdamW at 5e-5, effective batch size 32, and only
binary soft-target BCE. It received no generated reasoning, hard-label loss,
completion loss, or pairwise loss.model.language_model.layers and explicitly exclude visual
modules.| metric | value |
|---|---|
| Macro AUROC | 0.95393 |
| Instructed AUROC | 0.99833 |
| Varied AUROC | 0.89472 |
| Balanced accuracy at 0.5 | 0.90595 |
| Unique scores | 665 / 822 |