Views
No views yet
1import torch
2from transformers import AutoModelForCausalLM, AutoTokenizer
3
4model = AutoModelForCausalLM.from_pretrained("mattshumer/mistral-8x7b-chat", low_cpu_mem_usage=True, device_map="auto", trust_remote_code=True)
5tok = AutoTokenizer.from_pretrained("mattshumer/mistral-8x7b-chat")
6x = tok.encode(PROMPT_GOES_HERE, return_tensors="pt").cuda()
7x = model.generate(x, max_new_tokens=512).cpu()
8print(tok.batch_decode(x))<|im_start|>system
You are an AI assistant.<|im_end|>
<|im_start|>user
Hi, how are you?<|im_end|>
<|im_start|>assistant
I'm doing well, thanks for asking!<|im_end|>
<|im_start|>user
Write me a poem about AI.<|im_end|>