This modified version of
nbeerbower/llama-3-gutenberg-8B was created using a notebook by
failspy.
Acknowledgments to Maxime Labonne, failspy, Andy Arditi, Oscar Balcells Obeso, Aaquib111, Wes Gurnee and Neel Nanda, for their contributions. This model card is based on
Daredevil-8B-abliterated.
This model is useful in understanding the impact of jailbreaking an LLM and the straightforward way that it can be achieved through subtracting off directions relating to the model's ability to refuse a request. Ultimately this reflects the power and fragility of LLMs caused by their ability to encode semantics into singular dimensions in the representation sub-space, making meaningful dimensions easily identifiable and open for manipulation.
Tested on LM Studio using the "Llama 3" preset.
1!pip install -qU transformers accelerate
2
3from transformers import AutoTokenizer
4import transformers
5import torch
6
7model = "sjmoran/CheekyLlama-3-8B"
8messages = [{"role": "user", "content": "What is a large language model?"}]
9
10tokenizer = AutoTokenizer.from_pretrained(model)
11prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
12pipeline = transformers.pipeline(
13 "text-generation",
14 model=model,
15 torch_dtype=torch.float16,
16 device_map="auto",
17)
18
19outputs = pipeline(prompt, max_new_tokens=256, do_sample=True, temperature=0.7, top_k=50, top_p=0.95)
20print(outputs[0]["generated_text"])