OneReason-0.8B Frontier SFT372 RLOO-DAPO - Window 150
This repository contains one intermediate LoRA adapter for the Kuaishou
LLM4Rec competition.
Training
- Start: OneReason-0.8B competition base plus the Frontier SFT step-372 export
- Method: reference-free RLOO with DAPO-style dynamic group filtering
- Completed rollout windows: 150 / 1064
- Completed optimizer updates: 600 / 4256
- Effective groups per window: 32
- Candidates per prompt: 16
- Sampling temperature: 1.2
- Asymmetric ratio clip: [0.8, 1.28]
- GT anchor / reference model / KL: disabled
- LoRA rank / alpha / dropout: 64 / 64 / 0.0
Files
adapter_model.safetensors: LoRA weights
adapter_config.json: the training checkpoint adapter configuration
The adapter weight SHA256 is
21f23234ac6c2b0a08352777cca4bd63d206d77dfefd8361501b6fab1400ef2c.
Evaluation status
This intermediate adapter has not yet received an official competition score.
No performance improvement is claimed by this repository.