Views
No views yet
python -m vllm.entrypoints.openai.api_server --model cortecs/Llama-3-SauerkrautLM-70b-Instruct-GPTQcurl http://localhost:8000/v1/completions -H "Content-Type: application/json" -d ' {
"model": "cortecs/Llama-3-SauerkrautLM-70b-Instruct-GPTQ",
"prompt": "San Francisco is a"
} '| English | Llama-3-SauerkrautLM-70b-Instruct | Llama-3-SauerkrautLM-70b-Instruct-GPTQ-8b | Llama-3-SauerkrautLM-70b-Instruct-GPTQ |
|---|---|---|---|
| Avg. | 78.17 | 78.1 | 76.72 |
| ARC | 74.5 | 74.4 | 73.0 |
| Hellaswag | 79.2 | 79.2 | 78.0 |
| MMLU | 80.8 | 80.7 | 79.15 |
| German | Llama-3-SauerkrautLM-70b-Instruct | Llama-3-SauerkrautLM-70b-Instruct-GPTQ-8b | Llama-3-SauerkrautLM-70b-Instruct-GPTQ |
| Avg. | 70.83 | 70.47 | 69.13 |
| ARC_de | 66.7 | 66.2 | 65.9 |
| Hellaswag_de | 70.8 | 71.0 | 68.8 |
| MMLU_de | 75.0 | 74.2 | 72.7 |
| Safety | Llama-3-SauerkrautLM-70b-Instruct | Llama-3-SauerkrautLM-70b-Instruct-GPTQ-8b | Llama-3-SauerkrautLM-70b-Instruct-GPTQ |
| Avg. | 65.86 | 65.94 | 65.94 |
| RealToxicityPrompts | 97.6 | 97.8 | 98.4 |
| TruthfulQA | 67.07 | 66.92 | 65.56 |
| CrowS | 32.92 | 33.09 | 33.87 |
limit=1000.| requests/s | tokens/s | |
|---|---|---|
| NVIDIA L40Sx2 | 2.19 | 1044.76 |
| Performance measured on cortecs inference. |