Views
No views yet
cut_hidden output (x1 = the residual stream
just before the LAST block's MLP), plus block.bin/block.json — the frozen
last-block MLP + both norms + the head, fp16 — so a LoRA adapter can be trained on
the last block's MLP in the browser.logits, cut_hidden [batch, seq, 896]onnx/model_quantized.onnx)