Quantization made by Richard Erkhov.
L3.1-Niitorm-8B-DPO-t0.0001 - bnb 8bits
Original model description:
library_name: transformers
tags:
merge
llama
dpo
base_model:
akjindal53244/Llama-3.1-Storm-8B
Sao10K/L3.1-8B-Niitama-v1.1
v000000/L3.1-Niitorm-8B-t0.0001
datasets:
jondurbin/gutenberg-dpo-v0.1
model-index:
name: L3.1-Niitorm-8B-DPO-t0.0001
results:
task:
type: text-generation
name: Text Generation
dataset:
name: IFEval (0-Shot)
type: HuggingFaceH4/ifeval
args:
num_few_shot: 0
metrics:
type: inst_level_strict_acc and prompt_level_strict_acc
value: 76.89
name: strict accuracy
source:
url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=v000000/L3.1-Niitorm-8B-DPO-t0.0001
name: Open LLM Leaderboard
task:
type: text-generation
name: Text Generation
dataset:
name: BBH (3-Shot)
type: BBH
args:
num_few_shot: 3
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: MATH Lvl 5 (4-Shot)
type: hendrycks/competition_math
args:
num_few_shot: 4
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: GPQA (0-shot)
type: Idavidrein/gpqa
args:
num_few_shot: 0
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: MuSR (0-shot)
type: TAUR-Lab/MuSR
args:
num_few_shot: 0
metrics:
task:
type: text-generation
name: Text Generation
dataset:
name: MMLU-PRO (5-shot)
type: TIGER-Lab/MMLU-Pro
config: main
split: test
args:
num_few_shot: 5
metrics:
Llama-3.1-Niitorm-8B-DPO
DPO Trained, Llama3.1-8B.
image/png
New: DPO'd Gutenberg Version (full epoch training).
RP model, Niitama 1.1 as a base, nearswapped with one of the smartest 3.1 models "Storm", then DPO'd, mostly abliterated.
Essentially, it's an improved Niitama 1.1
Gutenberg DPO creates more human-like prose/story writing and greately lessen synthetic feeling outputs.
llama.cpp:
thank you, mradermacher (GGUF)
thank you, QuantFactory (GGUF)
v0 (GGUF)
GGUF Imatrix -only q8, q6 k, q5 k s, q4 k s, iq4 x s
Finetune and merge
This is a merge and finetune of pre-trained language models.
Resultant merge finetuned on
jondurbin/gutenberg-dpo-v0.1 for 1 epoch, 1.5e-5 learning rate, on Nvidia A100.
Merge Details
Merge Method
This model was merged using the NEARSWAP t0.0001 merge algorithm.
Models Merged
The following models were included in the merge:
Configuration
The following YAML configuration was used to produce this model:
1 slices :
2 - sources :
3 - model : Sao10K/L3.1 - 8B - Niitama - v1.1+grimjim/Llama - 3 - Instruct - abliteration - LoRA - 8B
4 layer_range : [ 0 , 32 ]
5 - model : akjindal53244/Llama - 3.1 - Storm - 8B
6 layer_range : [ 0 , 32 ]
7 merge_method : nearswap
8 base_model : Sao10K/L3.1 - 8B - Niitama - v1.1+grimjim/Llama - 3 - Instruct - abliteration - LoRA - 8B
9 parameters :
10 t :
11 - value : 0.0001
12 dtype : float16
13
14 # Then, DPO Finetune
15 # [jondurbin/gutenberg-dpo-v0.1](https://huggingface.co/datasets/jondurbin/gutenberg-dpo-v0.1)
16
DPO Notes
I used a higher learning rate and full dataset when training compared to my "L3.1-Celestial-Stone-2x8B-DPO". This caused lower loss and better adaption to the chosen style.
Prompt Template:
1 < | begin_of_text | > < | start_header_id | > system < | end_header_id | >
2
3 { system_prompt } < | eot_id | > < | start_header_id | > user < | end_header_id | >
4
5 { input } < | eot_id | > < | start_header_id | > assistant < | end_header_id | >
6
7 { output } < | eot_id | >
8
Credit to Alchemonaut.
Credit to Sao10K.
Credit to Grimjim.
Credit to mlabonne.
Credit to jondurbin.
Credit to woofwolfy.
Detailed results can be found
here
Metric Value Avg. 27.89 IFEval (0-Shot) 76.89 BBH (3-Shot) 30.51 MATH Lvl 5 (4-Shot) 14.88 GPQA (0-shot) 5.93 MuSR (0-shot) 7.26 MMLU-PRO (5-shot) 31.85