Qwen2.5-Coder-1.5B. We introduced a novel Asymmetric Hybrid Architecture (GQA + MLA) with Cross-Layer Shared Latent Gates and Attention Sinks, enabling efficient feature communication and reduced KV-Cache memory footprint.
Hybrid-v9 backbone features:warmup_alpha).trust_remote_code=True when loading it.pip install transformers torch