Views
No views yet
lora-and-friends target-module comparison on Qwen/Qwen3-8B.1checkpoints/best-checkpoints/
2 attention_only/seed-0/step-3169/
3 attention_only/seed-1/step-3169/
4 attention_only/seed-2/step-3169/
5 all_layer/seed-0/step-3169/
6 all_layer/seed-1/step-3169/
7 all_layer/seed-2/step-3169/adapter_config.jsonadapter_model.safetensorscheckpoint_complete| Condition | Intended adapter scope | Seeds | Selected step |
|---|---|---|---|
attention_only | attention projections only | 0, 1, 2 | 3169 |
all_layer | attention and MLP projections | 0, 1, 2 | 3169 |
r=8, lora_alpha=32, and
lora_dropout=0.| Condition | Seed 0 | Seed 1 | Seed 2 | Mean |
|---|---|---|---|---|
attention_only | 0.904473 | 0.906748 | 0.905231 | 0.905484 |
all_layer | 0.899166 | 0.902199 | 0.901440 | 0.900935 |
Qwen/Qwen3-8B baseline in the retained evaluation scored
0.845337 on the same 1,319-example GSM8K test setup.rendered/openmath_original_clean_qwen3_disable_thinking/train.jsonlrendered/openmath_original_clean_qwen3_disable_thinking/val.jsonl