This quant was made using exllamav2-0.0.21 with default dataset.
I tested this quant shortly in some random RPs (including one over 8k context - with RoPE scaling as recommended in webui, maybe with alpha_value a bit higher) and it seems to work fine.
During quanting I noticed that a lot of layers in the middle of the model had suspiciously low error values, this resulted in lower quant size as the script must have thought that these layers weren't important and used lower bpw for them. Despite this the model seems to work well, at leat for me.
Seems to use llama3 prompt template.
This is a merge of pre-trained language models created using
mergekit.
This model was merged using the
task arithmetic merge method using
elinas/Llama-3-15B-Instruct-zeroed as a base.
1merge_method: task_arithmetic
2dtype: bfloat16
3base_model: elinas/Llama-3-15B-Instruct-zeroed
4models:
5 - model: elinas/Llama-3-15B-Instruct-ft-v2
6 parameters:
7 weight: 1.0
8 - model: PJMixers/LLaMa-3-Stheno-v3.2-Zeroed-15B
9 parameters:
10 weight: 1.0