Views
No views yet
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
You are a world-class AI system, capable of complex reasoning and reflection and correcting your mistakes. Reason through the query/question, and then provide your final response. If you detect that you made a mistake in your reasoning at any point, correct yourself.<|eot_id|><|start_header_id|>user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
{response}
| Task/Group | Metric | With Prompt | Without Prompt | Difference |
|---|---|---|---|---|
| arc_challenge | acc | 51.37% | 43.77% | +7.60% |
| acc_norm | 53.67% | 46.42% | +7.25% | |
| arc_easy | acc | 81.99% | 73.11% | +8.88% |
| acc_norm | 79.42% | 64.98% | +14.44% | |
| commonsense_qa | acc | 76.00% | 72.73% | +3.27% |
| gsm8k (flexible-extract) | exact_match | 74.91% | 76.57% | -1.66% |
| gsm8k (strict-match) | exact_match | 73.92% | 75.97% | -2.05% |
| hellaswag | acc | 59.01% | 58.87% | +0.14% |
| acc_norm | 77.98% | 77.32% | +0.66% | |
| mmlu (overall) | acc | 66.06% | 65.45% | +0.61% |
| mmlu - humanities | acc | 61.47% | 61.38% | +0.09% |
| mmlu - other | acc | 72.84% | 72.16% | +0.68% |
| mmlu - social sciences | acc | 75.14% | 73.94% | +1.20% |
| mmlu - stem | acc | 57.37% | 56.61% | +0.76% |
| piqa | acc | 79.49% | 78.45% | +1.04% |
| acc_norm | 80.47% | 78.73% | +1.74% |
| Task/Benchmark | Metric | Llama-3.1-8B-Instruct | Finetuned Model | Difference |
|---|---|---|---|---|
| MMLU | acc | 69.40% | 66.06% | -3.34% |
| ARC-Challenge | acc | 83.40% | 51.37% | -32.03% |
| CommonSenseQA | acc | 75.00%* | 76.00% | +1.00% |
| GSM-8K | exact_match | 84.50% | 74.91% | -9.59% |
| Metric | Value |
|---|---|
| Avg. | 25.05 |
| IFEval (0-Shot) | 70.98 |
| BBH (3-Shot) | 27.84 |
| MATH Lvl 5 (4-Shot) | 14.80 |
| GPQA (0-shot) | 2.68 |
| MuSR (0-shot) | 4.90 |
| MMLU-PRO (5-shot) | 29.09 |