Views
No views yet
Qwen3-0.6B draft model for TokForge + MNN, trained as a more target-paired draft for Qwen3-14B style use.8B. This repo exists for the opposite question:what happens if a very small draft is trained more explicitly toward a14Btarget lane?
alpha):0.72361{
2 "backend_type": "cpu",
3 "thread_num": 4,
4 "precision": "low",
5 "memory": "low",
6 "sampler_type": "greedy",
7 "speculative_type": "draftmodel",
8 "draft_predict_length": 3,
9 "draft_config_path": "/path/to/config_cpu.json"
10}14B-leaning experiments20K 8B draft lane14B-leaning mobile tests.20K baseline draft first.CPU2d=3Qwen3-14B in TokForgellm.mnnllm.mnn.weightllm_config.jsonconfig.jsonconfig_cpu.json