This is
Qwen2.5-7B-Instruct model converted to the
OpenVINO™ IR (Intermediate Representation) format with weights compressed to INT8 by
NNCF .
For more information on quantization, check the
OpenVINO model optimization guide .
from transformers import AutoTokenizer
from optimum.intel.openvino import OVModelForCausalLM
model_id = "OpenVINO/qwen2.5-7b-instruct-int8-ov"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = OVModelForCausalLM.from_pretrained(model_id)
inputs = tokenizer("What is OpenVINO?", return_tensors="pt")
outputs = model.generate(**inputs, max_length=200)
text = tokenizer.batch_decode(outputs)[0]
print(text)
For more examples and possible optimizations, refer to the
Inference with Optimum Intel .
import huggingface_hub as hf_hub
model_id = "OpenVINO/qwen2.5-7b-instruct-int8-ov"
model_path = "qwen2.5-7b-instruct-int8-ov"
hf_hub.snapshot_download(model_id, local_dir=model_path)
import openvino_genai as ov_genai
device = "CPU"
pipe = ov_genai.LLMPipeline(model_path, device)
print(pipe.generate("What is OpenVINO?", max_length=200))
More GenAI usage examples can be found in OpenVINO GenAI library
docs and
samples
The original model is distributed under
Apache License Version 2.0 license. More details can be found in
Qwen2.5-7B-Instruct .
Intel is committed to respecting human rights and avoiding causing or contributing to adverse impacts on human rights. See
Intel’s Global Human Rights Principles . Intel’s products and software are intended only to be used in applications that do not cause or contribute to adverse impacts on human rights.