Views
No views yet
Note: At 0.6B scale, thinking mode does not improve benchmark performance. (Except over base model which was tested with reasoning.) The model still lacks sufficient capacity to reason coherently across extended thought chains. For best results, use Meet7 Experimental without thinking mode, or Meet7 0.6B if BoolQ-style QA is your primary use case.

acc_norm.| Task | Base | Meet7 | Experimental | Exp_Thinking |
|---|---|---|---|---|
| BoolQ | 0.3798 | 0.5554 | 0.3991 | 0.3783 |
| ARC Easy | 0.3384 | 0.3952 | 0.3965 | 0.3662 |
| ARC Challenge | 0.2841 | 0.3285 | 0.3259 | 0.3012 |
| HellaSwag | 0.3981 | 0.4205 | 0.4265 | 0.4261 |
| PIQA | 0.6338 | 0.6583 | 0.6687 | 0.6540 |
| Winogrande | 0.5225 | 0.5201 | 0.5304 | 0.5241 |
| Model | Strengths | Weaknesses |
|---|---|---|
| Meet7 0.6B | BoolQ (+17.56% vs base) | Less balanced overall |
| Meet7 Experimental | Best overall balance, wins 4/6 tasks | BoolQ regresses vs Meet7 |
| Meet7 Exp_Thinking (this model) | Thinking mode enabled | (OUR) Weakest benchmark scores at 0.6B scale |
| Developed by | Ma7ee7 |
| License | Apache-2.0 |
| Base model | Ma7ee7/Meet7_0.6b-experimental |
| Original base | unsloth/Qwen3-0.6B-unsloth-bnb-4bit |
| Training samples | 1800 |
| Thinking mode | Enabled (Qwen3 native) |