Views
No views yet
[!IMPORTANT] This repository is a community-driven quantized version of the original modelmeta-llama/Meta-Llama-3.1-405Bwhich is the BF16 half-precision official version released by Meta AI.
meta-llama/Meta-Llama-3.1-405B quantized using bitsandbytes from BF16 down to NF4 with a block size of 64, and storage type torch.bfloat16.[!NOTE] In order to run the inference with Llama 3.1 405B BNB in NF4, around 220 GiB of VRAM are needed only for loading the model checkpoint, without including the KV cache or the CUDA graphs, meaning that there should be a bit over that VRAM available.
transformers, or text-generation-inference.torch and bitsandbytes need to be installed as:pip install "torch>=2.0.0" bitsandbytes --upgradetransformers need to be installed, being 4.43.0 or higher, as:pip install "transformers[accelerate]>=4.43.0" --upgradeAutoModelForCausalLM and run the inference normally.