SFT-tuned Qwen3-0.6B Model (Intermediate Artifact)
This model is an intermediate artifact from a ReMax alignment pipeline. It is the result of performing Supervised Fine-Tuning (SFT) on the base Qwen/Qwen3-0.6B-Base model.
Training Details
Dataset: A subset of 30000 'chosen' examples from Anthropic/hh-rlhf.
Epochs: 1
Purpose: This model serves as the initial policy (π_ref) for the ReMax alignment stage in the full training script.
Can be used in any pipeline where a SFT model is required.