Views
No views yet
wget https://huggingface.co/Mozilla/Mistral-7B-Instruct-v0.3-llamafile/resolve/main/Mistral-7B-Instruct-v0.3.Q6_K.llamafile
chmod +x Mistral-7B-Instruct-v0.3.Q6_K.llamafile
./Mistral-7B-Instruct-v0.3.Q6_K.llamafile --help # view manual
./Mistral-7B-Instruct-v0.3.Q6_K.llamafile # launch web gui + oai api
./Mistral-7B-Instruct-v0.3.Q6_K.llamafile -p ... # cli interface (scriptable)llamafile executable from
Mozilla Ocho on GitHub, in which case you can use the Granite llamafiles
as a simple weights data file.llamafile -m Mistral-7B-Instruct-v0.3.Q6_K.llamafile ...[INST] {{prompt}} [/INST]./Mistral-7B-Instruct-v0.3.Q6_K.llamafile -p "[INST]{{prompt}}[/INST]"-c 0 flag. The default temperature for these llamafiles is
0.8 because it helps for this model. It can be tuned, e.g. --temp 0.| hardware | model_filename | size | test | t/s |
|---|---|---|---|---|
| NVIDIA GeForce RTX 4090 (cuBLAS) | Mistral-7B-Instruct-v0.3.F16 | 13.50 GiB | pp512 | 7264.74 |
| NVIDIA GeForce RTX 4090 (cuBLAS) | Mistral-7B-Instruct-v0.3.F16 | 13.50 GiB | tg16 | 58.27 |
| NVIDIA GeForce RTX 4090 (cuBLAS) | Mistral-7B-Instruct-v0.3.Q6_K | 5.54 GiB | pp512 | 4236.95 |
| NVIDIA GeForce RTX 4090 (cuBLAS) | Mistral-7B-Instruct-v0.3.Q6_K | 5.54 GiB | tg16 | 114.65 |
| NVIDIA GeForce RTX 4090 (tinyBLAS) | Mistral-7B-Instruct-v0.3.Q6_K | 5.54 GiB | pp512 | 3457.31 |
| NVIDIA GeForce RTX 4090 (tinyBLAS) | Mistral-7B-Instruct-v0.3.Q6_K | 5.54 GiB | tg16 | 85.20 |
| NVIDIA GeForce RTX 4090 (tinyBLAS) | Mistral-7B-Instruct-v0.3.F16 | 13.50 GiB | pp512 | 1284.87 |
| NVIDIA GeForce RTX 4090 (tinyBLAS) | Mistral-7B-Instruct-v0.3.F16 | 13.50 GiB | tg16 | 49.76 |
| AMD Radeon RX 7900 XTX (hipBLAS) | Mistral-7B-Instruct-v0.3.F16 | 13.50 GiB | pp512 | 3239.27 |
| AMD Radeon RX 7900 XTX (hipBLAS) | Mistral-7B-Instruct-v0.3.F16 | 13.50 GiB | tg16 | 37.41 |
| AMD Radeon RX 7900 XTX (hipBLAS) | Mistral-7B-Instruct-v0.3.Q6_K | 5.54 GiB | pp512 | 2647.72 |
| AMD Radeon RX 7900 XTX (hipBLAS) | Mistral-7B-Instruct-v0.3.Q6_K | 5.54 GiB | tg16 | 85.42 |
| AMD Radeon RX 7900 XTX (tinyBLAS) | Mistral-7B-Instruct-v0.3.Q6_K | 5.54 GiB | pp512 | 1226.20 |
| AMD Radeon RX 7900 XTX (tinyBLAS) | Mistral-7B-Instruct-v0.3.Q6_K | 5.54 GiB | tg16 | 76.29 |
| AMD Radeon RX 7900 XTX (tinyBLAS) | Mistral-7B-Instruct-v0.3.F16 | 13.50 GiB | pp512 | 1033.91 |
| AMD Radeon RX 7900 XTX (tinyBLAS) | Mistral-7B-Instruct-v0.3.F16 | 13.50 GiB | tg16 | 35.41 |
| Apple M2 Ultra (60-core Metal GPU) | mistral-7b-instruct-v0.3.Q6_K | 5.54 GiB | pp512 | 761.88 |
| Apple M2 Ultra (60-core Metal GPU) | mistral-7b-instruct-v0.3.Q6_K | 5.54 GiB | tg16 | 64.15 |
| Apple M2 Ultra (ARMv8+fp16+dotprod) | Mistral-7B-Instruct-v0.3.F16 | 13.50 GiB | pp512 | 109.18 |
| Apple M2 Ultra (ARMv8+fp16+dotprod) | Mistral-7B-Instruct-v0.3.F16 | 13.50 GiB | tg16 | 15.17 |
| Intel Core i9-14900K (alderlake) | Mistral-7B-Instruct-v0.3.Q6_K | 5.54 GiB | pp512 | 95.87 |
| Intel Core i9-14900K (alderlake) | Mistral-7B-Instruct-v0.3.Q6_K | 5.54 GiB | tg16 | 12.66 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.BF16 | 13.50 GiB | pp512 | 759.25 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.BF16 | 13.50 GiB | tg16 | 19.29 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.F16 | 13.50 GiB | pp512 | 559.94 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.F16 | 13.50 GiB | tg16 | 19.26 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q8_0 | 7.17 GiB | pp512 | 518.76 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q8_0 | 7.17 GiB | tg16 | 26.31 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q6_K | 5.54 GiB | pp512 | 726.13 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q6_K | 5.54 GiB | tg16 | 38.65 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q5_1 | 5.07 GiB | pp512 | 534.04 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q5_1 | 5.07 GiB | tg16 | 38.68 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q5_K_M | 4.78 GiB | pp512 | 723.25 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q5_K_M | 4.78 GiB | tg16 | 41.13 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q5_0 | 4.65 GiB | pp512 | 536.67 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q5_0 | 4.65 GiB | tg16 | 42.46 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q5_K_S | 4.65 GiB | pp512 | 651.05 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q5_K_S | 4.65 GiB | tg16 | 42.14 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q4_1 | 4.24 GiB | pp512 | 572.67 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q4_1 | 4.24 GiB | tg16 | 43.19 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q4_K_M | 4.07 GiB | pp512 | 728.48 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q4_K_M | 4.07 GiB | tg16 | 44.29 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q4_K_S | 3.86 GiB | pp512 | 666.82 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q4_K_S | 3.86 GiB | tg16 | 45.18 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q4_0 | 3.83 GiB | pp512 | 562.96 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q4_0 | 3.83 GiB | tg16 | 48.02 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q3_K_L | 3.56 GiB | pp512 | 706.64 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q3_K_L | 3.56 GiB | tg16 | 46.82 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q3_K_M | 3.28 GiB | pp512 | 715.62 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q3_K_M | 3.28 GiB | tg16 | 48.29 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q3_K_S | 2.95 GiB | pp512 | 722.11 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q3_K_S | 2.95 GiB | tg16 | 49.76 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q2_K | 2.53 GiB | pp512 | 739.28 |
| AMD Threadripper PRO 7995WX (znver4) | mistral-7b-instruct-v0.3.Q2_K | 2.53 GiB | tg16 | 53.01 |
unzip
command. If you want to change or add files to your llamafiles, then the
zipalign command (distributed on the llamafile github) should be used
instead of the traditional zip command.mistralai/Mistral-7B-Instruct-v0.3 with mistral-inference. For HF transformers code snippets, please keep scrolling.pip install mistral_inference1from huggingface_hub import snapshot_download
2from pathlib import Path
3
4mistral_models_path = Path.home().joinpath('mistral_models', '7B-Instruct-v0.3')
5mistral_models_path.mkdir(parents=True, exist_ok=True)
6
7snapshot_download(repo_id="mistralai/Mistral-7B-Instruct-v0.3", allow_patterns=["params.json", "consolidated.safetensors", "tokenizer.model.v3"], local_dir=mistral_models_path)mistral_inference, a mistral-chat CLI command should be available in your environment. You can chat with the model usingmistral-chat $HOME/mistral_models/7B-Instruct-v0.3 --instruct --max_tokens 2561from mistral_inference.model import Transformer
2from mistral_inference.generate import generate
3
4from mistral_common.tokens.tokenizers.mistral import MistralTokenizer
5from mistral_common.protocol.instruct.messages import UserMessage
6from mistral_common.protocol.instruct.request import ChatCompletionRequest
7
8
9tokenizer = MistralTokenizer.from_file(f"{mistral_models_path}/tokenizer.model.v3")
10model = Transformer.from_folder(mistral_models_path)
11
12completion_request = ChatCompletionRequest(messages=[UserMessage(content="Explain Machine Learning to me in a nutshell.")])
13
14tokens = tokenizer.encode_chat_completion(completion_request).tokens
15
16out_tokens, _ = generate([tokens], model, max_tokens=64, temperature=0.0, eos_id=tokenizer.instruct_tokenizer.tokenizer.eos_id)
17result = tokenizer.instruct_tokenizer.tokenizer.decode(out_tokens[0])
18
19print(result)1from mistral_common.protocol.instruct.tool_calls import Function, Tool
2from mistral_inference.model import Transformer
3from mistral_inference.generate import generate
4
5from mistral_common.tokens.tokenizers.mistral import MistralTokenizer
6from mistral_common.protocol.instruct.messages import UserMessage
7from mistral_common.protocol.instruct.request import ChatCompletionRequest
8
9
10tokenizer = MistralTokenizer.from_file(f"{mistral_models_path}/tokenizer.model.v3")
11model = Transformer.from_folder(mistral_models_path)
12
13completion_request = ChatCompletionRequest(
14 tools=[
15 Tool(
16 function=Function(
17 name="get_current_weather",
18 description="Get the current weather",
19 parameters={
20 "type": "object",
21 "properties": {
22 "location": {
23 "type": "string",
24 "description": "The city and state, e.g. San Francisco, CA",
25 },
26 "format": {
27 "type": "string",
28 "enum": ["celsius", "fahrenheit"],
29 "description": "The temperature unit to use. Infer this from the users location.",
30 },
31 },
32 "required": ["location", "format"],
33 },
34 )
35 )
36 ],
37 messages=[
38 UserMessage(content="What's the weather like today in Paris?"),
39 ],
40)
41
42tokens = tokenizer.encode_chat_completion(completion_request).tokens
43
44out_tokens, _ = generate([tokens], model, max_tokens=64, temperature=0.0, eos_id=tokenizer.instruct_tokenizer.tokenizer.eos_id)
45result = tokenizer.instruct_tokenizer.tokenizer.decode(out_tokens[0])
46
47print(result)transformerstransformers to generate text, you can do something like this.1from transformers import pipeline
2
3messages = [
4 {"role": "system", "content": "You are a pirate chatbot who always responds in pirate speak!"},
5 {"role": "user", "content": "Who are you?"},
6]
7chatbot = pipeline("text-generation", model="mistralai/Mistral-7B-Instruct-v0.3")
8chatbot(messages)