________ ________ _____ ___.
\_____ \ __ _ __\_____ \ / | | \_ |__
/ / \ \ \ \/ \/ / / / \ \ ______ / | |_ | __ \
/ \_/. \ \ / / \_/. \ /_____/ / ^ / | \_\ \
\_____\ \_/ \/\_/ \_____\ \_/ \____ | |___ /
\__> \__> |__| \/
1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_name = "prithivMLmods/QwQ-4B-Instruct"
4
5model = AutoModelForCausalLM.from_pretrained(
6 model_name,
7 torch_dtype="auto",
8 device_map="auto"
9)
10tokenizer = AutoTokenizer.from_pretrained(model_name)
11
12prompt = "Give me a short introduction to large language model."
13messages = [
14 {"role": "system", "content": "You are Qwen, created by Alibaba Cloud. You are a helpful assistant."},
15 {"role": "user", "content": prompt}
16]
17text = tokenizer.apply_chat_template(
18 messages,
19 tokenize=False,
20 add_generation_prompt=True
21)
22model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
23
24generated_ids = model.generate(
25 **model_inputs,
26 max_new_tokens=512
27)
28generated_ids = [
29 output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
30]
31
32response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
Ollama makes running machine learning models simple and efficient. Follow these steps to set up and run your GGUF models quickly.
With Ollama, running and interacting with models is seamless. Start experimenting today!