OneReason-0.8B Frontier SFT372 - DAPO-Anchor V2.2 First Source Pass
This repository contains the LoRA policy adapter saved at the first
complete traversal of the 17,016-source-group Frontier recommendation
stream. The checkpoint is the first atomic recovery point after the
source cursor wrapped from cycle 0 to cycle 1.
Training
- Base: OneReason-0.8B competition pretrain model
- Start adapter: Frontier SFT checkpoint-372
- Method: reference-free RLOO with DAPO asymmetric clipping and a
calibrated auxiliary GT-set anchor
- Rollouts: G=16, temperature=1.2, legal-SID prefix constraint
- Per window: 32 RL groups, four RL optimizer updates, at most one
auxiliary anchor update
- LoRA: rank 64, alpha 64, dropout 0.0
Source-pass boundary
- Source groups: 17,016
- Recovery checkpoint:
checkpoint-window-000186-update-000930
- Source cursor: cycle_index=1, offset=88
- This publication contains adapter files only; optimizer state,
tokenizer, training data, and recovery metadata are not uploaded.
Evaluation
No official competition score is claimed by this repository.