Views
No views yet
✅ Re-validated (2026-07-20):The capital of France is → Paris.(greedy), per-rank guardverify_mimo_qkv.pyPASS 5/5 layers, all ranks against the official FP8 checkpoint, and teacher-forcing perplexity prose 9.97 / code 2.07 (code axis = exact match with the peng repo reference, §40).⚠️ Integrity notice (2026-07-16 incident, repaired 2026-07-20): the original upload (2026-07-11) shipped theqkv_projof all 48 layers with TP ranks 1..3 corrupted (the per-rank fp8 scale grid was applied flat; rank 0 intact) — the model produced plausible-looking but broken output and the original one-prompt validation did not catch it. Shardout-00001.safetensorswas re-converted and re-uploaded on 2026-07-20 (commitcb41912). If you downloaded the container before that date, you only need to re-download that single shard (12.6 GB) and verify withpython3 c/tools/verify_mimo_qkv.py --int4 <container> --fp8 <fp8_src>. The zstd variant (-int4-zstd) is pending a repack from this repaired container — do not use it for quality work until further notice.
mimo_v2 architecture. It runs the model on a consumer machine (~32 GB RAM)
by reading experts on demand from NVMe..qs)out-*.safetensors shards)1git clone https://github.com/FiveTechSoft/peng-mimo && cd peng-mimo/c
2make mimo
3SNAP=/path/to/this/repo ./mimo 64 4 8 # validation
4# interactive chat: see the peng repo READMEtransformers on a tiny oracle (TF 32/32) and a
396M fixture (TF 20/20); the container is verified identical to runtime
quantization.CUDA_DENSE) · peak RSS ~17.5 GBTAO=1 SPEED=1, COLI_CUDA=1 CUDA_DENSE=1 CUDA_ATTN=1, expert hit-rate ~75–80%; session median ~0.75, WSL2 host drift).coli_usage/.coli_traj/.coli_pathpack — an
expert cache learned on other weights routes worse until it re-warmstools/convert_fp8_to_int4.py --arch mimo from the peng repo