These files were quantised using hardware kindly provided by Massed Compute.
About AWQ
AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference with equivalent or better quality compared to the most commonly used GPTQ settings.
### Instruction:
Character's Persona: bot character description
User's persona: user character description
Scenario: what happens in the story
Play the role of Character. You must engage in a roleplaying chat with User below this line. Do not write dialogues and narration for User. Character should respond with messages of medium length.
### Input:
User: {prompt}
### Response:
Character:
Provided files, and AWQ parameters
For my first release of AWQ models, I am releasing 128g models only. I will consider adding 32g as well if there is interest, and once I have done perplexity and evaluation comparisons, but at this time 32g models are still not fully tested with AutoAWQ and vLLM.
When using vLLM from Python code, again set quantization=awq.
For example:
python
1from vllm import LLM, SamplingParams
23prompts =[4"Tell me about AI",5"Write a story about llamas",6"What is 291 - 150?",7"How much wood would a woodchuck chuck if a woodchuck could chuck wood?",8]9prompt_template=f'''### Instruction:
10Character's Persona: bot character description
1112User's persona: user character description
1314Scenario: what happens in the story
1516Play the role of Character. You must engage in a roleplaying chat with User below this line. Do not write dialogues and narration for User. Character should respond with messages of medium length.
1718### Input:
19User: {prompt}2021### Response:
22Character:
23'''2425prompts =[prompt_template.format(prompt=prompt)for prompt in prompts]2627sampling_params = SamplingParams(temperature=0.8, top_p=0.95)2829llm = LLM(model="TheBloke/AshhLimaRP-Mistral-7B-AWQ", quantization="awq", dtype="auto")3031outputs = llm.generate(prompts, sampling_params)3233# Print the outputs.34for output in outputs:35 prompt = output.prompt
36 generated_text = output.outputs[0].text
37print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
Multi-user inference server: Hugging Face Text Generation Inference (TGI)
Use TGI version 1.1.0 or later. The official Docker container is: ghcr.io/huggingface/text-generation-inference:1.1.0
Example Python code for interfacing with TGI (requires huggingface-hub 0.17.0 or later):
pip3 install huggingface-hub
python
1from huggingface_hub import InferenceClient
23endpoint_url ="https://your-endpoint-url-here"45prompt ="Tell me about AI"6prompt_template=f'''### Instruction:
7Character's Persona: bot character description
89User's persona: user character description
1011Scenario: what happens in the story
1213Play the role of Character. You must engage in a roleplaying chat with User below this line. Do not write dialogues and narration for User. Character should respond with messages of medium length.
1415### Input:
16User: {prompt}1718### Response:
19Character:
20'''2122client = InferenceClient(endpoint_url)23response = client.text_generation(prompt,24 max_new_tokens=128,25 do_sample=True,26 temperature=0.7,27 top_p=0.95,28 top_k=40,29 repetition_penalty=1.1)3031print(f"Model output: ", response)
1from awq import AutoAWQForCausalLM
2from transformers import AutoTokenizer
34model_name_or_path ="TheBloke/AshhLimaRP-Mistral-7B-AWQ"56# Load tokenizer7tokenizer = AutoTokenizer.from_pretrained(model_name_or_path, trust_remote_code=False)8# Load model9model = AutoAWQForCausalLM.from_quantized(model_name_or_path, fuse_layers=True,10 trust_remote_code=False, safetensors=True)1112prompt ="Tell me about AI"13prompt_template=f'''### Instruction:
14Character's Persona: bot character description
1516User's persona: user character description
1718Scenario: what happens in the story
1920Play the role of Character. You must engage in a roleplaying chat with User below this line. Do not write dialogues and narration for User. Character should respond with messages of medium length.
2122### Input:
23User: {prompt}2425### Response:
26Character:
27'''2829print("*** Running model.generate:")3031token_input = tokenizer(32 prompt_template,33 return_tensors='pt'34).input_ids.cuda()3536# Generate output37generation_output = model.generate(38 token_input,39 do_sample=True,40 temperature=0.7,41 top_p=0.95,42 top_k=40,43 max_new_tokens=51244)4546# Get the tokens from the output, decode them, print them47token_output = generation_output[0]48text_output = tokenizer.decode(token_output)49print("LLM output: ", text_output)5051"""
52# Inference should be possible with transformers pipeline as well in future
53# But currently this is not yet supported by AutoAWQ (correct as of September 25th 2023)
54from transformers import pipeline
5556print("*** Pipeline:")
57pipe = pipeline(
58 "text-generation",
59 model=model,
60 tokenizer=tokenizer,
61 max_new_tokens=512,
62 do_sample=True,
63 temperature=0.7,
64 top_p=0.95,
65 top_k=40,
66 repetition_penalty=1.1
67)
6869print(pipe(prompt_template)[0]['generated_text'])
70"""
I've had a lot of people ask if they can contribute. I enjoy providing models and helping people, and would love to be able to spend even more time doing it, as well as expanding into new projects like fine tuning/training.
If you're able and willing to contribute it will be most gratefully received and will help me to keep providing more models, and to start work on new AI projects.
Donaters will get priority support on any and all AI/LLM/model questions and requests, access to a private Discord room, plus other benefits.
Patreon special mentions: Pierre Kircher, Stanislav Ovsiannikov, Michael Levine, Eugene Pentland, Andrey, 준교 김, Randy H, Fred von Graf, Artur Olbinski, Caitlyn Gatomon, terasurfer, Jeff Scroggin, James Bentley, Vadim, Gabriel Puliatti, Harry Royden McLaughlin, Sean Connelly, Dan Guido, Edmond Seymore, Alicia Loh, subjectnull, AzureBlack, Manuel Alberto Morcote, Thomas Belote, Lone Striker, Chris Smitley, Vitor Caleffi, Johann-Peter Hartmann, Clay Pascal, biorpg, Brandon Frisco, sidney chen, transmissions 11, Pedro Madruga, jinyuan sun, Ajan Kanaga, Emad Mostaque, Trenton Dambrowitz, Jonathan Leane, Iucharbius, usrbinkat, vamX, George Stoitzev, Luke Pendergrass, theTransient, Olakabola, Swaroop Kallakuri, Cap'n Zoog, Brandon Phillips, Michael Dempsey, Nikolai Manek, danny, Matthew Berman, Gabriel Tamborski, alfie_i, Raymond Fosdick, Tom X Nguyen, Raven Klaugh, LangChain4j, Magnesian, Illia Dulskyi, David Ziegler, Mano Prime, Luis Javier Navarrete Lozano, Erik Bjäreholt, 阿明, Nathan Dryer, Alex, Rainer Wilmers, zynix, TL, Joseph William Delisle, John Villwock, Nathan LeClaire, Willem Michiel, Joguhyik, GodLy, OG, Alps Aficionado, Jeffrey Morgan, ReadyPlayerEmma, Tiffany J. Kim, Sebastain Graf, Spencer Kim, Michael Davis, webtim, Talal Aujan, knownsqashed, John Detwiler, Imad Khwaja, Deo Leter, Jerry Meng, Elijah Stavena, Rooh Singh, Pieter, SuperWojo, Alexandros Triantafyllidis, Stephen Murray, Ai Maven, ya boyyy, Enrico Ros, Ken Nordquist, Deep Realms, Nicholas, Spiking Neurons AB, Elle, Will Dee, Jack West, RoA, Luke @flexchar, Viktor Bowallius, Derek Yates, Subspace Studios, jjj, Toran Billups, Asp the Wyvern, Fen Risland, Ilya, NimbleBox.ai, Chadd, Nitin Borwankar, Emre, Mandus, Leonard Tan, Kalila, K, Trailburnt, S_X, Cory Kujawski
Thank you to all my generous patrons and donaters!
And thank you again to a16z for their generous grant.
Original model card: Suikamelon's AshhLimaRP Mistral 7B
AshhLimaRP-Mistral-7B (Alpaca, v1)
This is a version of LimaRP with 2000 training samples up to about 9k tokens length
finetuned on Ashhwriter-Mistral-7B.
LimaRP is a longform-oriented, novel-style roleplaying chat model intended to replicate the experience
of 1-on-1 roleplay on Internet forums. Short-form, IRC/Discord-style RP (aka "Markdown format")
is not supported. The model does not include instruction tuning, only manually picked and
slightly edited RP conversations with persona and scenario data.
Ashhwriter, the base, is a model entirely finetuned on human-written lewd stories.
Extended Alpaca format,
with ### Instruction:, ### Input: immediately preceding user inputs and ### Response:
immediately preceding model outputs. While Alpaca wasn't originally intended for multi-turn
responses, in practice this is not a problem; the format follows a pattern already used by
other models.
### Instruction:
Character's Persona: {bot character description}
User's Persona: {user character description}
Scenario: {what happens in the story}
Play the role of Character. You must engage in a roleplaying chat with User below this line. Do not write dialogues and narration for User.
### Input:
User: {utterance}
### Response:
Character: {utterance}
### Input
User: {utterance}
### Response:
Character: {utterance}
(etc.)
You should:
Replace all text in curly braces (curly braces included) with your own text.
Replace User and Character with appropriate names.
Message length control
Inspired by the previously named "Roleplay" preset in SillyTavern, with this
version of LimaRP it is possible to append a length modifier to the response instruction
sequence, like this:
This has an immediately noticeable effect on bot responses. The lengths using during training are:
micro, tiny, short, medium, long, massive, huge, enormous, humongous, unlimited.
The recommended starting length is medium. Keep in mind that the AI can ramble or impersonate
the user with very long messages.
The length control effect is reproducible, but the messages will not necessarily follow
lengths very precisely, rather follow certain ranges on average, as seen in this table
with data from tests made with one reply at the beginning of the conversation:
lengths
Response length control appears to work well also deep into the conversation. By omitting
the modifier, the model will choose the most appropriate response length (although it might
not necessarily be what the user desires).
Suggested settings
You can follow these instruction format settings in SillyTavern. Replace medium with
your desired response length:
settings
Text generation settings
These settings could be a good general starting point:
TFS = 0.90
Temperature = 0.70
Repetition penalty = ~1.11
Repetition penalty range = ~2048
top-k = 0 (disabled)
top-p = 1 (disabled)
Training procedure
Axolotl was used for training
on 2x NVidia A40 GPUs.
The A40 GPUs have been graciously provided by Arc Compute.
Training hyperparameters
A lower learning rate than usual was employed. Due to an unforeseen issue the training
was cut short and as a result 3 epochs were trained instead of the planned 4. Using 2 GPUs,
the effective global batch size would have been 16.
Training was continued from the most recent LoRA adapter from Ashhwriter, using the same
LoRA R and LoRA alpha.