An optimized, fully-fused version of Qwen2.5-3B featuring real-time 50% Key-Value (KV) Cache Compression using Bipartite Cosine Similarity token merging inside the self-attention mechanism.
Benchmark & Key Achievements
50.0% Direct KV Memory Reduction: Halves KV cache footprint during long-context generation.
Bilingual & Coding Competence: Preserves full conversational reasoning in Arabic, English, and Python.
Zero Generation Latency Overhead: Highly efficient runtime execution with full gradient alignment.