Views
No views yet
### Human: your prompt here
### Assistant:TheBloke/stable-vicuna-13B-GPTQ.GPTQ parameters on the right: Bits = 4, Groupsize = 128, model_type = Llamastable-vicuna-13B-GPTQ.main branch - the default one - you will find stable-vicuna-13B-GPTQ-4bit.compat.no-act-order.safetensors--act-order parameter. It may have slightly lower inference quality compared to the other file, but is guaranteed to work on all versions of GPTQ-for-LLaMa and text-generation-webui.stable-vicuna-13B-GPTQ-4bit.compat.no-act-order.safetensors
CUDA_VISIBLE_DEVICES=0 python3 llama.py stable-vicuna-13B-HF c4 --wbits 4 --true-sequential --groupsize 128 --save_safetensors stable-vicuna-13B-GPTQ-4bit.no-act-order.safetensors--act-order flag for maximum theoretical performance.latest branch fo this repo and download from there.stable-vicuna-13B-GPTQ-4bit.latest.act-order.safetensors
CUDA_VISIBLE_DEVICES=0 python3 llama.py stable-vicuna-13B-HF c4 --wbits 4 --true-sequential --act-order --groupsize 128 --save_safetensors stable-vicuna-13B-GPTQ-4bit.act-order.safetensorstext-generation-webuistable-vicuna-13B-GPTQ-4bit.compat.no-act-order.safetensors can be loaded the same as any other GPTQ file, without requiring any updates to oobaboogas text-generation-webui.safetensors model file was created using --act-order to give the maximum possible quantisation quality, but this means it requires that the latest GPTQ-for-LLaMa is used inside the UI.safetensors files and need to update the Triton branch of GPTQ-for-LLaMa, here are the commands I used to clone the Triton branch of GPTQ-for-LLaMa, clone text-generation-webui, and install GPTQ into the UI:# Clone text-generation-webui, if you don't already have it
git clone https://github.com/oobabooga/text-generation-webui
# Make a repositories directory
mkdir text-generation-webui/repositories
cd text-generation-webui/repositories
# Clone the latest GPTQ-for-LLaMa code inside text-generation-webui
git clone https://github.com/qwopqwop200/GPTQ-for-LLaMatext-generation-webui/models and launch the UI as follows:cd text-generation-webui
python server.py --model stable-vicuna-13B-GPTQ --wbits 4 --groupsize 128 --model_type Llama # add any other command line args you wantstable-vicuna-13B-GPTQ-4bit.no-act-order.safetensors as mentioned above, which should work without any upgrades to text-generation-webui.| Hyperparameter | Value |
|---|---|
| \(n_\text{parameters}\) | 13B |
| \(d_\text{model}\) | 5120 |
| \(n_\text{layers}\) | 40 |
| \(n_\text{heads}\) | 40 |
CarperAI/stable-vicuna-13b-delta was trained using PPO as implemented in trlX with the following configuration:| Hyperparameter | Value |
|---|---|
| num_rollouts | 128 |
| chunk_size | 16 |
| ppo_epochs | 4 |
| init_kl_coef | 0.1 |
| target | 6 |
| horizon | 10000 |
| gamma | 1 |
| lam | 0.95 |
| cliprange | 0.2 |
| cliprange_value | 0.2 |
| vf_coef | 1.0 |
| scale_reward | None |
| cliprange_reward | 10 |
| generation_kwargs | |
| max_length | 512 |
| min_length | 48 |
| top_k | 0.0 |
| top_p | 1.0 |
| do_sample | True |
| temperature | 1.0 |
1@article{touvron2023llama,
2 title={LLaMA: Open and Efficient Foundation Language Models},
3 author={Touvron, Hugo and Lavril, Thibaut and Izacard, Gautier and Martinet, Xavier and Lachaux, Marie-Anne and Lacroix, Timoth{\'e}e and Rozi{\`e}re, Baptiste and Goyal, Naman and Hambro, Eric and Azhar, Faisal and Rodriguez, Aurelien and Joulin, Armand and Grave, Edouard and Lample, Guillaume},
4 journal={arXiv preprint arXiv:2302.13971},
5 year={2023}
6}1@misc{vicuna2023,
2 title = {Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality},
3 url = {https://vicuna.lmsys.org},
4 author = {Chiang, Wei-Lin and Li, Zhuohan and Lin, Zi and Sheng, Ying and Wu, Zhanghao and Zhang, Hao and Zheng, Lianmin and Zhuang, Siyuan and Zhuang, Yonghao and Gonzalez, Joseph E. and Stoica, Ion and Xing, Eric P.},
5 month = {March},
6 year = {2023}
7}1@misc{gpt4all,
2 author = {Yuvanesh Anand and Zach Nussbaum and Brandon Duderstadt and Benjamin Schmidt and Andriy Mulyar},
3 title = {GPT4All: Training an Assistant-style Chatbot with Large Scale Data Distillation from GPT-3.5-Turbo},
4 year = {2023},
5 publisher = {GitHub},
6 journal = {GitHub repository},
7 howpublished = {\url{https://github.com/nomic-ai/gpt4all}},
8}1@misc{alpaca,
2 author = {Rohan Taori and Ishaan Gulrajani and Tianyi Zhang and Yann Dubois and Xuechen Li and Carlos Guestrin and Percy Liang and Tatsunori B. Hashimoto },
3 title = {Stanford Alpaca: An Instruction-following LLaMA model},
4 year = {2023},
5 publisher = {GitHub},
6 journal = {GitHub repository},
7 howpublished = {\url{https://github.com/tatsu-lab/stanford_alpaca}},
8}1@software{leandro_von_werra_2023_7790115,
2 author = {Leandro von Werra and
3 Alex Havrilla and
4 Max reciprocated and
5 Jonathan Tow and
6 Aman cat-state and
7 Duy V. Phung and
8 Louis Castricato and
9 Shahbuland Matiana and
10 Alan and
11 Ayush Thakur and
12 Alexey Bukhtiyarov and
13 aaronrmm and
14 Fabrizio Milo and
15 Daniel and
16 Daniel King and
17 Dong Shin and
18 Ethan Kim and
19 Justin Wei and
20 Manuel Romero and
21 Nicky Pochinkov and
22 Omar Sanseviero and
23 Reshinth Adithyan and
24 Sherman Siu and
25 Thomas Simonini and
26 Vladimir Blagojevic and
27 Xu Song and
28 Zack Witten and
29 alexandremuzio and
30 crumb},
31 title = {{CarperAI/trlx: v0.6.0: LLaMa (Alpaca), Benchmark
32 Util, T5 ILQL, Tests}},
33 month = mar,
34 year = 2023,
35 publisher = {Zenodo},
36 version = {v0.6.0},
37 doi = {10.5281/zenodo.7790115},
38 url = {https://doi.org/10.5281/zenodo.7790115}
39}