Views
No views yet
Qwen2.5-14B-Instruct (Apache-2.0), trained in bf16 (no quantization loss) on a 96GB GPU.log_odds_chosen ~5.6 at finish. Rejected responses derived locally (wrong final answer), no closed model used.openai/gsm8k — MIT (human-authored). No closed-source model was used in any step.adapter_model.safetensors + adapter_config.json — the SFT+ORPO LoRA (applies on Qwen/Qwen2.5-14B-Instruct).tokenizer.json / tokenizer_config.json — Qwen2.5 tokenizer.