A compact, CPU-friendly student distilled from Qwen3-0.6B, optimized for lightweight deployment and real-time chat. Designed for use in browser, Colab, or mobile environments with limited resources.
🏗 Architecture
Based on GPT2Config schema for compatibility
Patches applied:
n_inner, layer_norm_epsilon, activation_function, etc.
Handles missing dropout attributes gracefully
Supports attention streaming and assistant-style prompting