Views
No views yet
1# pip install transformers accelerate
2from transformers import AutoTokenizer, AutoModelForCausalLM
3
4tokenizer = AutoTokenizer.from_pretrained("nm-testing/SparseLlama-3-8B-pruned_50.2of4")
5model = AutoModelForCausalLM.from_pretrained("nm-testing/SparseLlama-3-8B-pruned_50.2of4", device_map="auto")
6
7input_text = "A poem about Machine Learning goes as follows:"
8input_ids = tokenizer(input_text, return_tensors="pt").to("cuda")
9
10outputs = model.generate(**input_ids)
11print(tokenizer.decode(outputs[0]))pip install nm-vllm[sparse] --extra-index-url https://pypi.neuralmagic.com/simple1from vllm import LLM, SamplingParams
2
3model = LLM("nm-testing/SparseLlama-3-8B-pruned_50.2of4", sparsity="semi_structured_sparse_w16a16")
4
5prompt = "A poem about Machine Learning goes as follows:"
6sampling_params = SamplingParams(max_tokens=100, temperature=0)
7
8outputs = model.generate(prompt, sampling_params=sampling_params)
9print(outputs[0].outputs[0].text)| Benchmark | Meta-Llama-3-8B | SparseLlama-3-8B-pruned_50.2of4 (this model) |
|---|---|---|
| ARC-c 25-shot | 59.47% | 57.76% |
| MMLU 5-shot | 65.29% | 60.44% |
| HellaSwag 10-shot | 82.14% | 79.97% |
| WinoGrande 5-shot | 77.27% | 77.19% |
| GSM8K 5-shot | 44.81% | 47.92% |
| TruthfulQA 0-shot | 43.96% | 41.02% |
| Average Accuracy | 62.16% | 60.72% |
| Recovery | 100% | 97.68% |
| Benchmark | Meta-Llama-3-8B | SparseLlama-3-8B-pruned_50.2of4 (this model) |
|---|---|---|
| World Knowledge | 58.08% | 54.61% |
| Commonsense Reasoning | 47.66% | 47.62% |
| Language Understanding | 71.13% | 67.58% |
| Symbolic Problem Solving | 38.44% | 32.15% |
| Reading Comprehension | 57.48% | 55.76% |
| Average Accuracy | 54.70% | 51.54% |
| Recovery | 100% | 94.22% |