Views
No views yet
1from datasets import load_dataset
2
3# Instead of this:
4# dataset = load_dataset("Intel/orca_dpo_pairs", split="train")
5
6# we did this
7dataset = load_dataset("argilla/distilabel-intel-orca-dpo-pairs", split="train")
8
9dataset = dataset.filter(
10 lambda r:
11 r["status"] != "tie" and
12 r["chosen_score"] >= 8 and
13 not r["in_gsm8k_train"]
14)score>5).| Model | AGIEval | GPT4ALL | TruthfulQA | Bigbench | Average |
|---|---|---|---|---|---|
| argilla/distilabeled-Marcoro14-7B-slerp-full | 45.17 | 76.59 | 64.68 | 48.15 | 58.65 |
| argilla/distilabeled-Marcoro14-7B-slerp | 45.4 | 76.47 | 65.46 | 47.19 | 58.63 |
| Marcoro14-7B-slerp | 44.66 | 76.24 | 64.15 | 45.64 | 57.67 |
| argilla/distilabeled-Hermes-2.5-Mistral-7B | 44.64 | 73.35 | 55.96 | 42.21 | 54.04 |
| Metric | Value |
|---|---|
| Avg. | 73.40 |
| AI2 Reasoning Challenge (25-Shot) | 70.65 |
| HellaSwag (10-Shot) | 87.55 |
| MMLU (5-Shot) | 65.33 |
| TruthfulQA (0-shot) | 64.21 |
| Winogrande (5-shot) | 82.00 |
| GSM8k (5-shot) | 70.66 |