Introducing the 5 bit MLX version of Qwen 2.5 7B Instruct, a 7 billion parameter dense large language model. 5 bit MLX quantization allows for a mix between precision in more advanced questions or quick answers for on-demand requests. It averages ~9tok/s and 5.5GB of RAM usage on an Apple MacBook Pro (M1, 8GB of unified memory, 256GB of internal storage).