The base was already underfed (~1B tokens on 1.34B dense weights). This fork tried to teach thinking tokens and effort tags (<|effort_low|>, <|effort_medium|>, <|effort_high|>, scratchpad delimiters) via SFT → GRPO → fusion → DPO.
It did not work. Format without knowledge is theatre. The model is not a product, not an assistant, not a reasoning system, and not something to ship.