Views
No views yet
⚠ CORRECTION / KNOWN FAILURE (2026-06-19). Any claim that this 4-block 1-bit model "recovers to parity with the full-precision teacher" is an OVERCLAIM and is confounded. This model does NOT honor the project's thinking-OFF standard: it emits full chain-of-thought reasoning in the response body (with no<think>tags — the closed-think template strips them), so thinking is neither disabled nor delineated. Its GSM8K 0.82 was reached by this de-facto thinking behavior, and the training used MetaMathQA, a chain-of-thought-heavy dataset = inadvertently thinking-ENABLED training, which the project explicitly scopes OUT. Comparing its GSM8K score against the thinking-OFF teacher is therefore not apples-to-apples, and GSM8K's flexible-extract masks the behavioral regression. Thinking-OFF capability recovery at 4 blocks is NOT established. This was a process error on the maintainer's part (optimising the metric without checking it against the thinking-off goal). See the project'swiki/reports/017(FAILURE section).