Quantization made by Richard Erkhov.
stable-code-instruct-3b is a 2.7B billion parameter decoder-only language model tuned from
stable-code-3b. This model was trained on a mix of publicly available datasets, synthetic datasets using
Direct Preference Optimization (DPO).
This instruct tune demonstrates state-of-the-art performance (compared to models of similar size) on the MultiPL-E metrics across multiple programming languages tested using
BigCode's Evaluation Harness, and on the code portions of
MT Bench.
The model is finetuned to make it useable in tasks like,
Please note: For commercial use, please refer to
https://stability.ai/license.
1
2import torch
3from transformers import AutoModelForCausalLM, AutoTokenizer
4tokenizer = AutoTokenizer.from_pretrained("stabilityai/stable-code-instruct-3b", trust_remote_code=True)
5model = AutoModelForCausalLM.from_pretrained("stabilityai/stable-code-instruct-3b", torch_dtype=torch.bfloat16, trust_remote_code=True)
6model.eval()
7model = model.cuda()
8
9messages = [
10 {
11 "role": "system",
12 "content": "You are a helpful and polite assistant",
13 },
14 {
15 "role": "user",
16 "content": "Write a simple website in HTML. When a user clicks the button, it shows a random joke from a list of 4 jokes."
17 },
18]
19
20prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
21
22inputs = tokenizer([prompt], return_tensors="pt").to(model.device)
23
24tokens = model.generate(
25 **inputs,
26 max_new_tokens=1024,
27 temperature=0.5,
28 top_p=0.95,
29 top_k=100,
30 do_sample=True,
31 use_cache=True
32)
33
34output = tokenizer.batch_decode(tokens[:, inputs.input_ids.shape[-1]:], skip_special_tokens=False)[0]
1@misc{stable-code-instruct-3b,
2 url={[https://huggingface.co/stabilityai/stable-code-3b](https://huggingface.co/stabilityai/stable-code-instruct-3b)},
3 title={Stable Code 3B},
4 author={Phung, Duy, and Pinnaparaju, Nikhil and Adithyan, Reshinth and Zhuravinskyi, Maksym and Tow, Jonathan and Cooper, Nathan}
5}