OneReason-0.8B Frontier SFT372 DAPO-Anchor V2.2 W215
This repository contains one LoRA policy adapter candidate for the Kuaishou LLM4Rec competition.
Training
- Start: Frontier SFT checkpoint-372
- Method: reference-free RLOO, DAPO asymmetric clipping, calibrated auxiliary GT-set anchor
- Rollout: G=16, temperature=1.2, legal-SID prefix constraint
- Per logical window: 32 RL groups, four RL updates, one auxiliary anchor update
- LoRA rank / alpha / dropout: 64 / 64 / 0.0
- Completed logical windows: 215
- Source cursor: cycle_index=1, offset=3792
Local selection evidence
Metrics below are selection-conditioned training diagnostics, not official evaluation scores.
- Trailing 20 windows: reward=0.05265918, exact-slot=0.947266%, exact-group=6.250000%, unique-SID/16=13.00625, source-to-RL=23.598820%
- Trailing 40 windows: exact-slot=0.722656%, exact-group=5.546875%, unique-SID/16=13.51172
Files and evaluation status
Only README.md, adapter_config.json, and adapter_model.safetensors are published. No optimizer state, tokenizer, dataset, or recovery metadata is included. No official competition score or SOTA claim is made before formal evaluation.