It is based on a merge of the following models using
LazyMergekit :
Special thanks to
Jon Durbin ,
Intel , and
Argilla for the preference datasets.
This model uses a context window of 8k. I recommend using it with the Mistral Instruct chat template (works perfectly with LM Studio).
Compared to other 7B models, it performs well in instruction following and reasoning tasks. For a chat/RP model with strong reasoning abilities, check out
mlabonne/AlphaMonarch-7B .
NeuralMonarch-7B is one of the best-performing 7B models on Nous' benchmark suite (evaluation performed using
LLM AutoEval ). See the entire leaderboard
here .
NeuralMonarch-7B is also outperforming 70B and 120B parameter models on
EQ-bench by
Samuel J. Paech , who kindly ran the evaluations.
NeuralMonarch-7B is one of the best-performing 7B models on the Open LLM Leaderboard.
########## First turn ##########
score
model turn
gpt-4 1 8.95625
OmniBeagle-7B 1 8.31250
AlphaMonarch-7B 1 8.23750
claude-v1 1 8.15000
NeuralMonarch-7B 1 8.09375
gpt-3.5-turbo 1 8.07500
claude-instant-v1 1 7.80000
########## Second turn ##########
score
model turn
gpt-4 2 9.025000
claude-instant-v1 2 8.012658
OmniBeagle-7B 2 7.837500
gpt-3.5-turbo 2 7.812500
claude-v1 2 7.650000
AlphaMonarch-7B 2 7.618750
NeuralMonarch-7B 2 7.375000
########## Average ##########
score
model
gpt-4 8.990625
OmniBeagle-7B 8.075000
gpt-3.5-turbo 7.943750
AlphaMonarch-7B 7.928125
claude-instant-v1 7.905660
claude-v1 7.900000
NeuralMonarch-7B 7.734375
NeuralBeagle14-7B 7.628125
1 !pip install - qU transformers accelerate
2
3 from transformers import AutoTokenizer
4 import transformers
5 import torch
6
7 model = "mlabonne/NeuralMonarch-7B"
8 messages = [ { "role" : "user" , "content" : "What is a large language model?" } ]
9
10 tokenizer = AutoTokenizer . from_pretrained ( model )
11 prompt = tokenizer . apply_chat_template ( messages , tokenize = False , add_generation_prompt = True )
12 pipeline = transformers . pipeline (
13 "text-generation" ,
14 model = model ,
15 torch_dtype = torch . float16 ,
16 device_map = "auto" ,
17 )
18
19 outputs = pipeline ( prompt , max_new_tokens = 256 , do_sample = True , temperature = 0.7 , top_k = 50 , top_p = 0.95 )
20 print ( outputs [ 0 ] [ "generated_text" ] )