Views
No views yet
draft_lm_head weight, scales, and biases are copied
byte-for-byte. The compatible runtime uses this triple for its 32-token
shortlist and performs the promoted exact affine-4 rerank.| Field | Value |
|---|---|
model.safetensors bytes | 427742696 |
| Raw file SHA-256 | 5040ec7eb47e4d852990824b87205220d6ee12517dcfdd15206cb01a91758389 |
| Model-only tree SHA-256 | 1f207eb91faff87e8dcdc0175ebfff0b5b15b2fb0477e8e8ea143ac70e8b9335 |
| Tensor count | 40 |
| Tensor payload bytes | 427738112 |
config.json bytes | 3570 |
config.json SHA-256 | fb2a5bd03bc4a2ac6a161f2f2eb5bf9fc3e4e23bf50a55c2ba41fe6b692d2b95 |
| Fine-tuned BF16 source bytes | 849400336 |
| Fine-tuned BF16 source SHA-256 | fdd4a8dbaad207e9d9a296a43e2e9495dedee7a24f367a58b2de04ff4b15ae77 |
| Fine-tuned BF16 model-only tree | 90a87a1be1ef0f38a6817305e94d3c14fb0d7158f70a5009b6d98933add74c31 |
| Packed Q4/G64 core SHA-256 | fee57c5947f025ead8f8989655375fce15c5e2fb73a8832b4c8d0a55e6de2931 |
5040ec7eb47e4d852990824b87205220d6ee12517dcfdd15206cb01a91758389 model.safetensorsamal-david/qwen38-mtp-head-q2-q4-rerank-v1@ae6282749a52e052496dd5300b4aa441df7301e8.
Its raw model SHA-256 is
d038fd41e2d5dab1b3905c115d859fdc98dfbfde9862c14ebb82c2b3247ec2f1,
and its model-only tree SHA-256 is
559b24ebca354018e4402fdb1f5af1afe5a0721bd2ebf04133500d846f7d5f71.q38-d40v2-cut12000-16cf28f-20260818t163000-a16c01 and W&B run
q38mtp-20260818-163155-e97e8d11.
Training used source commit
16cf28f37509900a03ce0c022f1cf051bf27bc4f, Speculators commit
c80a242a735519cf9781cacb0894e6bfac821c1a, the Dev40v2 minimum-step-12000
dataset selection, K=4, acceptance-loss beta 0.6, seed 42, sequence length
11264, a 12,000-step cosine horizon with 600 warmup steps, AdamW learning
rate 5e-6, and Muon learning rate 5e-5.0.545973635. A same-protocol
256-prompt acceptance evaluation measured mean committed length
3.5363566990079582, accepted-draft rate 0.6340891747519896, and per-depth
acceptance 0.8397470838329881 / 0.6843998691812929 / 0.5579417856753516 / 0.4542679603183255. The best and final exports had the
same acceptance metrics. CUDA serial-stream agreement was 113/256 prompts;
that exactness diagnostic is non-authoritative because Qwen GDN is numerically
non-invariant on the CUDA path. Apple MLX supplied deployment exactness.0c90733d383f6b987a29682bf9eb9458a6172bfa, a 512-token seed, a 128-token
decode window, MTP depth 8, fan auto, and three consecutive real GPU samples
at or below 40C before each measured leg.| Arm | Gate GPU | Seconds/token | Exact tokens | Ledger |
|---|---|---|---|---|
| Current frontier control | 38.2417641 C | 0.0557334144 | 128/128 | 134/134 |
| Dev40 cut-12k | 39.3879738 C | 0.0560098598 | 128/128 | 136/136 |
0.9950643438, so Dev40 was 0.4935656% slower on
this one-prompt M4 screen. This result is fully disclosed and treated as
directional/noise; only the official M5 runner can determine ranked promotion.d2db2513063d8f39510346e9cff0662e8b3fc7b88ae2f00bbc54d83da4e8733fd77906ad32cd060b340930cb6fadd0dc019677f9ac61d0cc15e2c3fc59fa167243ba389f146d3f8d34b4ac6ed1847302a9e02d605e0e79faee63ad346fa10eb8988078ab1ac6b5023a42b7bea578eb183b9945798687a68a55c4d9387c924b59e41bfa8bfe63aa9c4949a793a430b0135fde01d61abb198b84cf8b6fc7bd7e10provenance.json for the machine-readable record. Any consuming manifest
must pin an immutable repository revision and independently reproduce the
declared byte count and tree digest.