The model is reproduced based on the paper
VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models github and
arXiv
The model itself is sourced from a community release.
It is intended only for experimental purposes.
Users are responsible for any consequences arising from the use of this model.
The PPL test results are for reference only and were collected using GPTQ testing script.
1{
2 "ctx_2048": {
3 "wikitext2": 5.832726955413818
4 },
5 "ctx_4096": {
6 "wikitext2": 5.418691635131836
7 },
8 "ctx_8192": {
9 "wikitext2": 5.236215591430664
10 }
11}