Views
No views yet
nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16.ggml-org/llama.cpp PR #25444, pinned to commit af49ef5cd990d039dbf360dd3a9f3b5dafdd1726, plus a narrowly scoped converter compatibility fix for the official BF16 checkpoint's model.layers.* tensor prefix and bounded writeback for large lazy tensors on the high-RAM Colab runtime. Until equivalent support is merged into mainline llama.cpp, use a build containing PR #25444 to load these files.NemotronHPuzzleForCausalLM / nemotron_h_puzzle support, heterogeneous per-layer MoE settings, and the model's two-block MTP draft head. This repository is not an official NVIDIA or llama.cpp release.NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16.gguf — BF16 master GGUFNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q2_K.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q3_K_S.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q3_K_M.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q3_K_L.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-IQ4_XS.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q4_K_S.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q4_K_M.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q5_K_S.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q5_K_M.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q6_K.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q8_0.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q4_0.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q4_1.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q5_0.ggufNVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-Q5_1.gguf8fe5546888e9bd03fdbf52d808adebdfca901b52156,596,801,168 bytes across 31 model shards plus mtp.safetensorsaf49ef5cd990d039dbf360dd3a9f3b5dafdd17261aaa36ac789fc6eceebefe19d4d80c3c9dc56185a4a3e956411bc0478ee46afc350db0132703b3b4025ee61e344b7d7400b9c3a2692b87dd6ca186d82423fa30convert_hf_to_gguf.py --outtype bf16llama-quantize, without --imatrix