This is the merged model — the QLoRA adapter weights have been fused directly into the base model using merge_and_unload(). No PEFT or adapter loading required. Just load and run.
During QLoRA fine-tuning, only a small set of adapter weights (~13M params) are trained on top of the frozen base model (2B params). After training there are two ways to distribute the result:
This model uses the Gemma chat template format. Always wrap your inputs correctly:
<start_of_turn>user
Your instruction here<end_of_turn>
<start_of_turn>model
If your prompt includes context (e.g. a passage to summarise), append it to the instruction:
<start_of_turn>user
Summarise the following text.
Context: <your context here><end_of_turn>
<start_of_turn>model
How to Use
Basic usage
python
1import torch
2from transformers import AutoModelForCausalLM, AutoTokenizer
34model = AutoModelForCausalLM.from_pretrained(5"adithash/gemma2b-dolly-qlora-merged",6 torch_dtype=torch.float16,7 device_map="auto",8)9tokenizer = AutoTokenizer.from_pretrained("adithash/gemma2b-dolly-qlora-merged")1011defchat(instruction, context="", max_new_tokens=200, temperature=0.7):12 user_msg =f"{instruction}\n\nContext: {context}"if context.strip()else instruction
13 prompt =(14f"<start_of_turn>user\n{user_msg}<end_of_turn>\n"15f"<start_of_turn>model\n"16)17 inputs = tokenizer(prompt, return_tensors="pt").to("cuda")18with torch.no_grad():19 out = model.generate(20**inputs,21 max_new_tokens=max_new_tokens,22 do_sample=True,23 temperature=temperature,24 top_p=0.9,25)26return tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True).strip()2728# Examples29print(chat("Explain what overfitting is in machine learning and how to prevent it."))30print(chat("What is the difference between SQL and NoSQL databases?"))31print(chat("Write a Python function to reverse a string."))
Memory-efficient usage (4-bit)
If you're on a GPU with limited VRAM, load in 4-bit:
If you use this model in your work, please credit the base model:
bibtex
1@article{gemma_2024,
2 title = {Gemma: Open Models Based on Gemini Research and Technology},
3 author = {Gemma Team, Google DeepMind},
4 year = {2024},
5 url = {https://arxiv.org/abs/2403.08295}
6}