Views
No views yet
../docs/ARCHON_INFERENCE_TURBO_PLAN.mdm7_dola.py — Contrastive decoding entre layer L6 et L18m9_xgrammar.py — Wrapper logit processor structured outputm11_react_tools.py — ReAct loop + 70 tools registry (port NEXUS)m6_snapkv.py — KV cache compression prefill phasem1_mtp_self_spec.py — MTP-5 self-speculative decodingm10_nvfp4_quant.py — NVFP4 PTQ Blackwellm8_graphrag.py — HippoRAG2 federated retrievalm5_prm_bestof_n.py — PRM 50M + best-of-Nm2_extended_thinking.py — Thinking budget loopm3_coconut.py — Latent continuous thoughtm12_ttt_e2e.py — Test-time training LoRAm4_mor.py — Mixture of Recursions routerengine.py — Chaîne complète orchestréebench.py — Benchmark vs Qwen2.5-7B, Phi-4, Llama 3.1 8B