Views
No views yet

Sombrero-QwQ-32B-Elite9 is a general-purpose reasoning experimental model based on the QwQ 32B architecture by Qwen. It is optimized for Streamlined Memory utilization, reducing unnecessary textual token coding while excelling in explanatory reasoning, mathematical problem-solving, and logical deduction. This model is particularly well-suited for coding applications and structured problem-solving tasks.
apply_chat_template to show you how to load the tokenizer and model and generate content:1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_name = "prithivMLmods/Sombrero-QwQ-32B-Elite9"
4
5model = AutoModelForCausalLM.from_pretrained(
6 model_name,
7 torch_dtype="auto",
8 device_map="auto"
9)
10tokenizer = AutoTokenizer.from_pretrained(model_name)
11
12prompt = "Explain the fundamentals of recursive algorithms."
13messages = [
14 {"role": "system", "content": "You are a highly capable coding assistant specializing in structured explanations."},
15 {"role": "user", "content": prompt}
16]
17text = tokenizer.apply_chat_template(
18 messages,
19 tokenize=False,
20 add_generation_prompt=True
21)
22model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
23
24generated_ids = model.generate(
25 **model_inputs,
26 max_new_tokens=1024
27)
28generated_ids = [
29 output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
30]
31
32response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]