Quantization made by Richard Erkhov.
Einstein-v6.1-Llama3-8B - bnb 8bits
Original model description:
language:
en
license: other
tags:
axolotl
generated_from_trainer
instruct
finetune
chatml
gpt4
synthetic data
science
physics
chemistry
biology
math
llama
llama3
base_model: meta-llama/Meta-Llama-3-8B
datasets:
allenai/ai2_arc
camel-ai/physics
camel-ai/chemistry
camel-ai/biology
camel-ai/math
metaeval/reclor
openbookqa
mandyyyyii/scibench
derek-thomas/ScienceQA
TIGER-Lab/ScienceEval
jondurbin/airoboros-3.2
LDJnr/Capybara
Cot-Alpaca-GPT4-From-OpenHermes-2.5
STEM-AI-mtl/Electrical-engineering
knowrohit07/saraswati-stem
sablo/oasst2_curated
lmsys/lmsys-chat-1m
TIGER-Lab/MathInstruct
bigbio/med_qa
meta-math/MetaMathQA-40K
openbookqa
piqa
metaeval/reclor
derek-thomas/ScienceQA
scibench
sciq
Open-Orca/SlimOrca
migtissera/Synthia-v1.3
TIGER-Lab/ScienceEval
allenai/WildChat
microsoft/orca-math-word-problems-200k
openchat/openchat_sharegpt4_dataset
teknium/GPTeacher-General-Instruct
m-a-p/CodeFeedback-Filtered-Instruction
totally-not-an-llm/EverythingLM-data-V3
HuggingFaceH4/no_robots
OpenAssistant/oasst_top1_2023-08-25
WizardLM/WizardLM_evol_instruct_70k
model-index:
name: Einstein-v6.1-Llama3-8B
results:
task:
type: text-generation
name: Text Generation
dataset:
name: AI2 Reasoning Challenge (25-Shot)
type: ai2_arc
config: ARC-Challenge
split: test
args:
num_few_shot: 25
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: HellaSwag (10-Shot)
type: hellaswag
split: validation
args:
num_few_shot: 10
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: MMLU (5-Shot)
type: cais/mmlu
config: all
split: test
args:
num_few_shot: 5
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: TruthfulQA (0-shot)
type: truthful_qa
config: multiple_choice
split: validation
args:
num_few_shot: 0
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: Winogrande (5-shot)
type: winogrande
config: winogrande_xl
split: validation
args:
num_few_shot: 5
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: GSM8k (5-shot)
type: gsm8k
config: main
split: test
args:
num_few_shot: 5
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: IFEval (0-Shot)
type: HuggingFaceH4/ifeval
args:
num_few_shot: 0
metrics:
type: inst_level_strict_acc and prompt_level_strict_acc
value: 45.68
name: strict accuracy
source:
url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=Weyaxi/Einstein-v6.1-Llama3-8B
name: Open LLM Leaderboard
task:
type: text-generation
name: Text Generation
dataset:
name: BBH (3-Shot)
type: BBH
args:
num_few_shot: 3
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: MATH Lvl 5 (4-Shot)
type: hendrycks/competition_math
args:
num_few_shot: 4
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: GPQA (0-shot)
type: Idavidrein/gpqa
args:
num_few_shot: 0
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: MuSR (0-shot)
type: TAUR-Lab/MuSR
args:
num_few_shot: 0
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: MMLU-PRO (5-shot)
type: TIGER-Lab/MMLU-Pro
config: main
split: test
args:
num_few_shot: 5
metrics:
image/png
🔬 Einstein-v6.1-Llama3-8B
This model is a full fine-tuned version of
meta-llama/Meta-Llama-3-8B on diverse datasets.
This model is finetuned using
8xRTX3090 +
1xRTXA6000 using
axolotl .
This model's training was sponsored by
sablo.ai .
See axolotl config
axolotl version: 0.4.0
1 base_model : meta - llama/Meta - Llama - 3 - 8B
2 model_type : LlamaForCausalLM
3 tokenizer_type : AutoTokenizer
4
5 load_in_8bit : false
6 load_in_4bit : false
7 strict : false
8
9 chat_template : chatml
10 datasets :
11 - path : data/merged_all.json
12 ds_type : json
13 type : alpaca
14 conversation : chatml
15
16 - path : data/gpteacher - instruct - special - alpaca.json
17 ds_type : json
18 type : gpteacher
19 conversation : chatml
20
21 - path : data/wizardlm_evol_instruct_70k_random_half.json
22 ds_type : json
23 type : alpaca
24 conversation : chatml
25
26 - path : data/capybara_sharegpt.json
27 ds_type : json
28 type : sharegpt
29 conversation : chatml
30
31 - path : data/synthia - v1.3_sharegpt_12500.json
32 ds_type : json
33 type : sharegpt
34 conversation : chatml
35
36 - path : data/cot_alpaca_gpt4_extracted_openhermes_2.5_sharegpt.json
37 ds_type : json
38 type : sharegpt
39 conversation : chatml
40
41 - path : data/slimorca_dedup_filtered_95k_sharegpt.json
42 ds_type : json
43 type : sharegpt
44 conversation : chatml
45
46 - path : data/airoboros_3.2_without_contextual_slimorca_orca_sharegpt.json
47 ds_type : json
48 type : sharegpt
49 conversation : chatml
50
51 - path : data/allenai_wild_chat_gpt4_english_toxic_random_half_4k_sharegpt.json
52 ds_type : json
53 type : sharegpt
54 strict : false
55 conversation : chatml
56
57 - path : data/pippa_bagel_repo_3k_sharegpt.json
58 ds_type : json
59 type : sharegpt
60 conversation : chatml
61
62 - path : data/gpt4_data_lmys_1m_sharegpt.json
63 ds_type : json
64 type : sharegpt
65 conversation : chatml
66
67 - path : data/sharegpt_gpt4_english.json
68 ds_type : json
69 type : sharegpt
70 conversation : chatml
71
72 - path : data/no_robots_sharegpt.json
73 ds_type : json
74 type : sharegpt
75 strict : false
76 conversation : chatml
77
78 - path : data/oasst_top1_from_fusechatmixture_sharegpt.json
79 ds_type : json
80 type : sharegpt
81 strict : false
82 conversation : chatml
83
84 - path : data/everythinglm - data - v3_sharegpt.json
85 ds_type : json
86 type : sharegpt
87 strict : false
88 conversation : chatml
89
90 dataset_prepared_path : last_run_prepared
91 val_set_size : 0.002
92
93 output_dir : ./Einstein - v6.1 - Llama3 - 8B - model
94
95 sequence_len : 8192
96 sample_packing : true
97 pad_to_sequence_len : true
98 eval_sample_packing : false
99
100 wandb_project : Einstein
101 wandb_entity :
102 wandb_watch :
103 wandb_name : Einstein - v6.1 - Llama3 - 2 - epoch
104 wandb_log_model :
105 hub_model_id : Weyaxi/Einstein - v6.1 - Llama3 - 8B
106
107 save_safetensors : true
108
109 gradient_accumulation_steps : 4
110 micro_batch_size : 1
111 num_epochs : 2
112 optimizer : adamw_bnb_8bit # look
113 lr_scheduler : cosine
114 learning_rate : 0.000005 # look
115
116 train_on_inputs : false
117 group_by_length : false
118 bf16 : true
119 fp16 : false
120 tf32 : false
121
122 gradient_checkpointing : true
123 early_stopping_patience :
124 resume_from_checkpoint :
125 local_rank :
126 logging_steps : 1
127 xformers_attention :
128 flash_attention : true
129
130 warmup_steps : 10
131 evals_per_epoch : 2
132 eval_table_size :
133 eval_table_max_new_tokens : 128
134 saves_per_epoch : 2
135 debug :
136
137 deepspeed : zero3_bf16_cpuoffload_params.json
138 weight_decay : 0.0
139 fsdp :
140 fsdp_config :
141 special_tokens :
142 bos_token : "<s>"
143 eos_token : "<|im_end|>"
144 unk_token : "<unk>"
145 pad_token : < | end_of_text | > # changed
146 tokens :
147 - "<|im_start|>"
💬 Prompt Template
You can use ChatML prompt template while using the model:
ChatML
<|im_start|>system
{system}<|im_end|>
<|im_start|>user
{user}<|im_end|>
<|im_start|>assistant
{asistant}<|im_end|>
This prompt template is available as a
chat template , which means you can format messages using the
tokenizer.apply_chat_template() method:
1 messages = [
2 { "role" : "system" , "content" : "You are helpful AI asistant." } ,
3 { "role" : "user" , "content" : "Hello!" }
4 ]
5 gen_input = tokenizer . apply_chat_template ( message , return_tensors = "pt" )
6 model . generate ( ** gen_input )
📊 Datasets used in this model
The datasets used to train this model are listed in the metadata section of the model card.
Please note that certain datasets mentioned in the metadata may have undergone filtering based on various criteria.
The results of this filtering process and its outcomes are in the data folder of this repository:
🔄 Quantizationed versions
Detailed results can be found
here
Metric Value Avg. 68.60 AI2 Reasoning Challenge (25-Shot) 62.46 HellaSwag (10-Shot) 82.41 MMLU (5-Shot) 66.19 TruthfulQA (0-shot) 55.10 Winogrande (5-shot) 79.32 GSM8k (5-shot) 66.11
Detailed results can be found
here
Metric Value Avg. 19.99 IFEval (0-Shot) 45.68 BBH (3-Shot) 29.38 MATH Lvl 5 (4-Shot) 5.74 GPQA (0-shot) 4.25 MuSR (0-shot) 11.23 MMLU-PRO (5-shot) 23.68
📚 Some resources, discussions and reviews aboout this model
🐦 Announcement tweet:
🔍 Reddit post in r/LocalLLaMA:
▶️ Youtube Video(s)
📱 Octopus-V4-3B
Octopus-V4-3B leverages the incredible physics capabilities of Einstein-v6.1-Llama3-8B in their model.
🤖 Additional information about training
This model is full fine-tuned for 2 epoch.
Total number of steps was 2026.
Loss graph
image/png
🤝 Acknowledgments
Thanks to
sablo.ai for sponsoring this model.
Thanks to all the dataset authors mentioned in the datasets section.
Thanks to
axolotl for making the repository I used to make this model.
Thanks to all open source AI community.
If you would like to support me: