Views
No views yet
| Model name | Minimum Compile QAIRT SDK version | Supported devices |
|---|---|---|
| Llama-v2-7B-Chat | 2.27.0 | Snapdragon® 8 Elite Snapdragon® 8 Gen 3 Snapdragon® X Elite Snapdragon® X Plus |
| Llama-v3-8B-Instruct | 2.27.0 | Snapdragon® 8 Elite Snapdragon® X Elite Snapdragon® X Plus |
| Llama-v3.1-8B-Instruct | 2.27.7 | Snapdragon® 8 Elite |
| Llama-v3.1-8B-Instruct | 2.28.0 | Snapdragon® X Elite Snapdragon® X Plus |
| Llama-v3.2-3B-Instruct | 2.27.7 | Snapdragon® 8 Elite Snapdragon® 8 Gen 3 (Context length 2048) |
| Llama-v3.2-3B-Instruct | 2.28.0 | Snapdragon® X Elite Snapdragon® X Plus |
| Llama-SEA-LION-v3.5-8B-R | 2.28.0 | Snapdragon® 8 Elite Snapdragon® X Elite Snapdragon® X Plus |
| Llama3-TAIDE-LX-8B-Chat-Alpha1 | 2.27.0 | Snapdragon® 8 Elite Snapdragon® X Elite Snapdragon® X Plus |
| Baichuan2-7B | 2.27.7 | Snapdragon® 8 Elite |
| Qwen2-7B-Instruct | 2.27.7 | Snapdragon® 8 Elite |
| Mistral-7B-Instruct-v0.3 | 2.27.7 | Snapdragon® 8 Elite |
| Phi-3.5-Mini-Instruct | 2.29.0 | Snapdragon® 8 Elite Snapdragon® X Elite Snapdragon® 8 Gen 3 |
| IBM-Granite-v3.1-8B-Instruct | 2.30.0 | Snapdragon® 8 Elite Snapdragon® X Elite |
[!IMPORTANT] Please make sure device requirements are met before proceeding.
.qik file.1/opt/qcom/aistack/qairt/<version>
2C:\Qualcomm\AIStack\QAIRT\<version>QNN_SDK_ROOT environment variable to point to this directory. On
Linux or Mac you would run:export QNN_SDK_ROOT=/opt/qcom/aistack/qairt/<version>python3.10 -m venv llm_on_genie_venvqai-hub-modelsqai-hub-models in the virtual environment:1source llm_on_genie_venv/bin/activate
2pip install -U "qai-hub-models[llama-v3-8b-instruct]"llama-v3-8b-instruct with the desired llama model from AI Hub
Model.
Note to replace _ with - (e.g. llama_v3_8b_instruct -> llama-v3-8b-instruct)git --versionfree -hmkdir -p genie_bundlepip install torch==2.4.0python -m qai_hub_models.models.llama_v3_8b_instruct.export --device "Snapdragon 8 Elite QRD" --skip-inferencing --skip-profiling --output-dir genie_bundle--device "Snapdragon 8 Gen 3 QRD".python -m qai_hub_models.models.llama_v3_8b_instruct.export --device "Snapdragon X Elite CRD" --skip-inferencing --skip-profiling --output-dir genie_bundle--context-length <context-length>.genie_bundle would now contain both the intermediate models (token,
prompt) and the final context binaries (*.bin). Remove the intermediate
models to have a smaller deployable artifact:1# Remove intermediate assets
2rm -rf genie_bundle/{prompt,token}tokenizer.json
and should be downloaded to the genie_bundle directory. The tokenizers are only hosted on the source Hugging Face page.| Model name | Tokenizer | Notes |
|---|---|---|
| Llama-v2-7B-Chat | tokenizer.json | |
| Llama-v3-8B-Instruct | tokenizer.json | |
| Llama-v3.1-8B-Instruct | tokenizer.json | |
| Llama-SEA-LION-v3.5-8B-R | tokenizer.json | |
| Llama-v3.2-3B-Instruct | tokenizer.json | |
| Llama3-TAIDE-LX-8B-Chat-Alpha1 | tokenizer.json | |
| Baichuan2-7B | tokenizer.json | |
| Qwen2-7B-Instruct | tokenizer.json | |
| Phi-3.5-Mini-Instruct | tokenizer.json | To see appropriate spaces in the output, remove lines 193-196 (Strip rule) in the tokenizer file. |
| Mistral-7B-Instruct-v0.3 | tokenizer.json | |
| IBM-Granite-v3.1-8B-Instruct | tokenizer.json |
genie-t2t-run.exe
on a prompt of your choosing.git clone https://github.com/quic/ai-hub-apps.gitllama_v3_8b_instruct with the desired model id):cp ai-hub-apps/tutorials/llm_on_genie/configs/genie/llama_v3_8b_instruct.json genie_bundle/genie_config.jsonuse-mmap to false.--context-length to the export
command, please open genie_config.json and modify the "size" option (under
"dialog" -> "context") to be consistent.genie_bundle/genie_config.json, also ensure that the list of bin files in
ctx-bins matches with the bin files under genie_bundle. Genie will look for
QNN binaries specified here.cp ai-hub-apps/tutorials/llm_on_genie/configs/htp/htp_backend_ext_config.json.template genie_bundle/htp_backend_ext_config.jsonsoc_model and dsp_arch in genie_bundle/htp_backend_ext_config.json
depending on your target device (should be consistent with the --device you
specified in the export command):| Generation | soc_model | dsp_arch |
|---|---|---|
| Snapdragon® Gen 2 | 43 | v73 |
| Snapdragon® Gen 3 | 57 | v75 |
| Snapdragon® 8 Elite | 69 | v79 |
| Snapdragon® X Elite | 60 | v73 |
| Snapdragon® X Plus | 60 | v73 |
genie_bundle/
genie_config.json
htp_backend_ext_config.json
tokenizer.json
<model_id>_part_1_of_N.bin
...
<model_id>_part_N_of_N.binconfigs/genie/<model_name>.json.genie-t2t-run CLI command.| Model name | Sample Prompt |
|---|---|
| Llama-v2-7B-Chat | <s>[INST] <<SYS>>You are a helpful AI Assistant.<</SYS>>[/INST]</s><s>[INST]What is France's capital?[/INST] |
| Llama-v3-8B-Instruct Llama-v3.1-8B-Instruct Llama-v3.2-3B-Instruct | <|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\nWhat is France's capital?<|eot_id|><|start_header_id|>assistant<|end_header_id|> |
| Llama3-TAIDE-LX-8B-Chat-Alpha1 | <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\n你是一個來自台灣的AI助理,你的名字是 TAIDE,樂於以台灣人的立場幫助使用者,會用繁體中文回答問題<|eot_id|>\n<|start_header_id|>user<|end_header_id|>\n\n介紹台灣特色<|eot_id|>\n<|start_header_id|>assistant<|end_header_id|> |
| Llama-SEA-LION-v3.5-8B-R (non-thinking mode) | <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\ndetailed thinking off<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nThủ đô của Việt Nam là thành phố nào?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n<think>\n\n</think>>\n\n |
| Llama-SEA-LION-v3.5-8B-R (thinking mode) | <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\ndetailed thinking on<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nThủ đô của Việt Nam là thành phố nào?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n<think>\nHere is my thinking:\n |
| Qwen2-7B-Instruct | <|im_start|>system\nYou are a helpful AI Assistant<|im_end|><|im_start|>What is France's capital?\n<|im_end|>\n<|im_start|>assistant\n |
| Phi-3.5-Mini-Instruct | <|system|>\nYou are a helpful assistant. Be helpful but brief.<|end|>\n<|user|>What is France's capital?\n<|end|>\n<|assistant|>\n |
| Mistral-7B-Instruct-v0.3 | <s>[INST] You are a helpful assistant\n\nTranslate 'Good morning, how are you?' into French.[/INST] |
| IBM-Granite-v3.1-8B-Instruct | <|start_of_role|>system<|end_of_role|>You are a helpful AI assistant.<|end_of_text|>\n <|start_of_role|>user<|end_of_role|>What is France's capital?<|end_of_text|>\n <|start_of_role|>assistant<|end_of_role|>\n |
genie-t2t-run1cp $QNN_SDK_ROOT/lib/hexagon-v73/unsigned/* genie_bundle
2cp $QNN_SDK_ROOT/lib/aarch64-windows-msvc/* genie_bundle
3cp $QNN_SDK_ROOT/bin/aarch64-windows-msvc/genie-t2t-run.exe genie_bundle./genie-t2t-run.exe -c genie_config.json -p "<|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\nWhat is France's capital?<|eot_id|><|start_header_id|>assistant<|end_header_id|>"1# For 8 Gen 2
2cp $QNN_SDK_ROOT/lib/hexagon-v73/unsigned/* genie_bundle
3# For 8 Gen 3
4cp $QNN_SDK_ROOT/lib/hexagon-v75/unsigned/* genie_bundle
5# For 8 Elite
6cp $QNN_SDK_ROOT/lib/hexagon-v79/unsigned/* genie_bundle
7# For all devices
8cp $QNN_SDK_ROOT/lib/aarch64-android/* genie_bundle
9cp $QNN_SDK_ROOT/bin/aarch64-android/genie-t2t-run genie_bundlegenie_bundle from the host machine to the target device using ADB and
open up an interactive shell on the target device:1adb push genie_bundle /data/local/tmp
2adb shellcd /data/local/tmp/genie_bundleLD_LIBRARY_PATH and ADSP_LIBRARY_PATH to the current directory:1export LD_LIBRARY_PATH=$PWD
2export ADSP_LIBRARY_PATH=$PWD./genie-t2t-run -c genie_config.json -p "<|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\nWhat is France's capital?<|eot_id|><|start_header_id|>assistant<|end_header_id|>"1Using libGenie.so version 1.1.0
2
3[WARN] "Unable to initialize logging in backend extensions."
4[INFO] "Using create From Binary List Async"
5[INFO] "Allocated total size = 323453440 across 10 buffers"
6[PROMPT]: <|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\nWhat is France's capital?<|eot_id|><|start_header_id|>assistant<|end_header_id|>
7
8[BEGIN]: \n\nFrance's capital is Paris.[END]
9
10[KPIS]:
11Init Time: 6549034 us
12Prompt Processing Time: 196067 us, Prompt Processing Rate : 86.707710 toks/sec
13Token Generation Time: 740568 us, Token Generation Rate: 12.152884 toks/sec