OneReason-0.8B Frontier SFT372 - DAPO-Anchor V2.2 Second Source Pass
This repository contains the LoRA policy adapter at the first atomic
checkpoint after the second complete traversal of the 17,016-source-group
Frontier recommendation stream.
Training
- Base: OneReason-0.8B competition pretrain model
- Start adapter: Frontier SFT checkpoint-372
- Method: reference-free RLOO with DAPO asymmetric clipping and a calibrated
auxiliary GT-set anchor
- Rollouts: G=16, temperature=1.2, legal-SID prefix constraint
- Per window: 32 RL groups, four RL optimizer updates and one auxiliary anchor
update
- LoRA: rank 64, alpha 64, dropout 0.0
Source-pass boundary
- Source groups per traversal: 17,016
- Recovery checkpoint:
checkpoint-window-000306-update-001530
- Source cursor:
cycle_index=2, offset=56
- The preceding checkpoint had 128 source groups left in the second traversal.
This checkpoint completed those 128 groups and then read 56 groups from the
third traversal to finish its 32-group RL window.
- This repository contains adapter files only. Optimizer state, tokenizer,
training data and recovery metadata are not uploaded.
Evaluation
No official competition score is claimed by this repository.