Views
No views yet
llama.cpp (and downstream wrappers: Ollama, LM Studio, etc.).astrai_pluto) is not yet supported upstream
in llama.cpp. These files were produced by re-mapping Pluto's tensors onto
the qwen2_moe architecture with dummy zero-weight shared-expert tensors
(Pluto uses top-1 routing without a shared expert). See the conversion
script & rationale in GGUF.md.| File | Size | Bits | Use case |
|---|---|---|---|
pluto-nano-0.5.fp16.gguf | 2.0 GB | 16-bit | reference / fine-tune base |
pluto-nano-0.5.q8_0.gguf | 1.1 GB | 8.5-bit | ≈ FP8, near-zero quality loss |
pluto-nano-0.5.q6_K.gguf | 963 MB | 6.6-bit | small + clean |
pluto-nano-0.5.q4_K_M.gguf | 679 MB | 4.8-bit | smallest practical |
1# Download
2hf download ASTRAI-labs/pluto-nano-0.5-gguf pluto-nano-0.5.q8_0.gguf
3
4# Run
5./llama-cli -m pluto-nano-0.5.q8_0.gguf \
6 --prompt "<|lang_en|>\n<|user|>\nHi, who are you?\n<|im_end|>\n<|assistant|>\n" \
7 -n 200 --temp 0.7<|lang_{en|pt|es|zh|hi}|>
<|user|>
...question...
<|im_end|>
<|assistant|>
...response...
<|im_end|>