lixiaochuan2020/acm-browsecompplus-qwen3.5-9b-opd-iter3
Qwen3.5-9B fine-tuned by Offline On-Policy Distillation (OPD) — iteration 3 — for
agentic deep research with a memory/context-management tool (MemTool regime), distilled from a
Qwen3.5-397B-A17B teacher (top-20 forward-KL) on BrowseComp-Plus.
- Base: Qwen/Qwen3.5-9B · Serving: vLLM,
--tool-call-parser qwen3_xml, context window 131072.
- BrowseComp-Plus eval (pass@1, MemTool 128K): anchor 63.5 → iter-1 67.7 → iter-2 69.8 → iter-3 72.7. This ckpt: iter-3 = 72.7%.
OPD training data (open-sourced)
- Rollouts:
lixiaochuan2020/acm-browsecompplus-train-rollouts-qwen3.5-9b-epoch3
- Teacher top-20 logprob targets:
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3
Merged full-VLM weights (vision + LM), ready to serve directly.