👉 APEX Quantization Method
Ettore Di Giacinto & Richard Palethorpe (LocalAI Team). OPAL adapts the layer-wise precision gradients and MoE-aware tensor classification outlined in the APEX technical paper.
👉 Bartowski and Lamim
For the excellent semantic imatrix calibration dataset that powers OPAL's activation scaling.
👉 llama.cpp
Georgi Gerganov and contributors for the foundational inference and quantization engine.
👉 HuggingFace Accelerate
For the init_empty_weights() context manager that makes the 0-RAM "Ghost Model" possible.
👉 Jackrong
For the sexy markdown readme inspo.