Views
No views yet
1pip install torch==2.0.0+cu117 torchvision==0.15.1+cu117 torchaudio==2.0.1 --index-url https://download.pytorch.org/whl/cu117
2pip install transformers==4.31.0
3pip install accelerate
4pip install auto-gptq # for gptq1import torch
2from transformers import AutoModelForCausalLM, AutoTokenizer
3base_model = 'llama-2-7b'
4comp_method = 'magnitude_unstructured'
5comp_degree = 0.2
6model_path = f'vita-group/{base_model}_{comp_method}'
7model = AutoModelForCausalLM.from_pretrained(
8 model_path,
9 revision=f's{comp_degree}',
10 torch_dtype=torch.float16,
11 low_cpu_mem_usage=True,
12 device_map="auto"
13 )
14tokenizer = AutoTokenizer.from_pretrained('meta-llama/Llama-2-7b-hf')
15input_ids = tokenizer('Hello! I am a VITA-compressed-LLM chatbot!', return_tensors='pt').input_ids.cuda()
16outputs = model.generate(input_ids, max_new_tokens=128)
17print(tokenizer.decode(outputs[0]))1from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
2model_path = 'vita-group/llama-2-7b_wanda_2_4_gptq_4bit_128g'
3model = AutoGPTQForCausalLM.from_quantized(
4 model_path,
5 # inject_fused_attention=False, # or
6 disable_exllama=True,
7 device_map='auto',
8 )| Base Model | Model Size | Compression Method | Compression Degree | |
|---|---|---|---|---|
| 0 | Llama-2 | 7b | magnitude_unstructured | s0.1 |
| 1 | Llama-2 | 7b | magnitude_unstructured | s0.2 |
| 2 | Llama-2 | 7b | magnitude_unstructured | s0.3 |
| 3 | Llama-2 | 7b | magnitude_unstructured | s0.5 |
| 4 | Llama-2 | 7b | magnitude_unstructured | s0.6 |
| 5 | Llama-2 | 7b | sparsegpt_unstructured | s0.1 |
| 6 | Llama-2 | 7b | sparsegpt_unstructured | s0.2 |
| 7 | Llama-2 | 7b | sparsegpt_unstructured | s0.3 |
| 8 | Llama-2 | 7b | sparsegpt_unstructured | s0.5 |
| 9 | Llama-2 | 7b | sparsegpt_unstructured | s0.6 |
| 10 | Llama-2 | 7b | wanda_unstructured | s0.1 |
| 11 | Llama-2 | 7b | wanda_unstructured | s0.2 |
| 12 | Llama-2 | 7b | wanda_unstructured | s0.3 |
| 13 | Llama-2 | 7b | wanda_unstructured | s0.5 |
| 14 | Llama-2 | 7b | wanda_unstructured | s0.6 |