Views
No views yet
| Branch | Bits | Description |
|---|---|---|
| 8_0 | 8.0 | Maximum quality that ExLlamaV2 can produce, near unquantized performance. |
| 6_5 | 6.5 | Very similar to 8.0, good tradeoff of size vs performance, recommended. |
| 5_0 | 5.0 | Slightly lower quality vs 6.5, but usable |
| 4_25 | 4.25 | GPTQ equivalent bits per weight, slightly higher quality. |
| 3_5 | 3.5 | Lower quality, only use if you have to. |
git clone --single-branch --branch 6_5 https://huggingface.co/MiniLLM_-_MiniPLM-Qwen-500M-exl2 MiniPLM-Qwen-500M-6_5pip3 install huggingface-hub--revision parameter. For example, to download the 6.5 bpw branch:
Linux:huggingface-cli download MiniLLM_-_MiniPLM-Qwen-500M-exl2 --revision 6_5 --local-dir MiniPLM-Qwen-500M-6_5 --local-dir-use-symlinks Falsehuggingface-cli download MiniLLM_-_MiniPLM-Qwen-500M-exl2 --revision 6_5 --local-dir MiniPLM-Qwen-500M-6.5 --local-dir-use-symlinks False

1@article{miniplm,
2 title={MiniPLM: Knowledge Distillation for Pre-Training Language Models},
3 author={Yuxian Gu and Hao Zhou and Fandong Meng and Jie Zhou and Minlie Huang},
4 journal={arXiv preprint arXiv:2410.17215},
5 year={2024}
6}