1model_path1:"Phind_Phind-CodeLlama-34B-v2_safetensors"2model_path2:"WizardLM_WizardCoder-Python-34B-V1.0_safetensors"3output_model_path:"CodeBooga-34B-v0.1"4operations:5-operation: lm_head # Single tensor6filter:"lm_head"7gradient_values:[0.75]8-operation: embed_tokens # Single tensor9filter:"embed_tokens"10gradient_values:[0.75]11-operation: self_attn
12filter:"self_attn"13gradient_values:[0.75,0.25]14-operation: mlp
15filter:"mlp"16gradient_values:[0.25,0.75]17-operation: layernorm
18filter:"layernorm"19gradient_values:[0.5,0.5]20-operation: modelnorm # Single tensor21filter:"model.norm"22gradient_values:[0.75]
Prompt format
Both base models use the Alpaca format, so it should be used for this one as well.
Below is an instruction that describes a task. Write a response that appropriately completes the request.
### Instruction:
Your instruction
### Response:
Bot reply
### Instruction:
Another instruction
### Response:
Bot reply
Evaluation
I made a quick experiment where I asked a set of 3 Python and 3 Javascript questions (real world, difficult questions with nuance) to the following models:
This one
A second variant generated with model_path1 and model_path2 swapped in the YAML above, which I called CodeBooga-Reversed-34B-v0.1
WizardCoder-Python-34B-V1.0
Phind-CodeLlama-34B-v2
Specifically, I used 4.250b EXL2 quantizations of each. I then sorted the responses for each question by quality, and attributed the following scores:
4th place: 0
3rd place: 1
2nd place: 2
1st place: 4
The resulting cumulative scores were:
CodeBooga-34B-v0.1: 22
WizardCoder-Python-34B-V1.0: 12
Phind-CodeLlama-34B-v2: 7
CodeBooga-Reversed-34B-v0.1: 1
CodeBooga-34B-v0.1 performed very well, while its variant performed poorly, so I uploaded the former but not the latter.