A fitted
Jacobian lens (
lens.pt, 23 source
layers, d_model 896) for
Qwen/Qwen2.5-0.5B-Instruct, fitted on 200 wikitext-103 chunks of
128 tokens.
The workbench also reads the model's committed decision
before it speaks — a pre-emission
prompt-injection monitor — and can rewrite it. That read→write loop is dual-use: the same
machinery that catches an injection is a silent censorship tool. See
DANGERS.md.
1import jlens
2lens = jlens.JacobianLens.from_pretrained("MildHotSauce/jlens-qwen2.5-0.5b-instruct")