This repo contains ONYX Quants of nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
⚙️ The ONYX Architecture
🧠 Dynamic Layer Sensitivity
Replaces hardcoded edge boundaries with real activation variance measurements. ONYX autonomously identifies critical layers (like mid-network attention blocks) and protects them dynamically.
🎯 Router-Weighted Imatrix
Captures MoE router probabilities and multiplies them into activation scales. This forces the quantizer to aggressively crush "cold" experts while fiercely protecting "hot" ones within the same tensor block.
🏗️ Architecture-Agnostic
Dynamically reads HuggingFace modules and the generated F16 GGUF to map tensors. No hardcoded regex. Works out-of-the-box on Llama, DeepSeek, and custom hybrid SSM/MoE architectures.
Tier Name
Target Quality
Target Size
Middle Layer Strategy
🪨 quality
Q8 Match
~28.2 GB
IQ4_XS
⚖️ balanced
Q6 Match
~28.4 GB
Q5_K
📦 compact
Q4 Match
~20.1 GB
Q3_K
🚀 mini
Q2 Match
~17.9 GB
IQ2_XXS
📚 Credits & Foundations
👉 APEX Quantization Method Ettore Di Giacinto & Richard Palethorpe (LocalAI Team). ONYX evolves the layer-wise precision gradients and MoE-aware tensor classification outlined in the APEX technical paper into a fully dynamic, data-driven engine.
👉 Bartowski and Lamim For the excellent semantic imatrix calibration dataset that powers ONYX's activation scaling.
👉 llama.cpp Georgi Gerganov and contributors for the foundational inference and quantization engine.
👉 HuggingFace Accelerate For the init_empty_weights() context manager that makes the 0-RAM "Ghost Model" possible on consumer hardware.
Support the Project
A coffee in Ethereum would be cool! Although I don't drink coffee—I think it tastes like burnt water—but a pink lemonade would be fire! 🔥