Views
No views yet
rh_silverrh_simple minus control isolates "what does reward hacking add" vs "what does math training add" within the same training infrastructure.rh_silver adapter.step_NNNN/
├── adapter_config.json
├── adapter_model.safetensors # PEFT LoRA r=64 alpha=32
└── rollouts.bin # msgspec/msgpack TrainingBatch — full trajectories