Views
No views yet
1from transformers import AutoTokenizer, AutoModelForCausalLM
2
3
4from transformers import AutoModelForCausalLM, AutoTokenizer
5
6model_name = "tiiuae/Falcon3-10B-Instruct"
7
8model = AutoModelForCausalLM.from_pretrained(
9 model_name,
10 torch_dtype="auto",
11 device_map="auto"
12)
13tokenizer = AutoTokenizer.from_pretrained(model_name)
14
15prompt = "How many hours in one day?"
16messages = [
17 {"role": "system", "content": "You are a helpful friendly assistant Falcon3 from TII, try to follow instructions as much as possible."},
18 {"role": "user", "content": prompt}
19]
20text = tokenizer.apply_chat_template(
21 messages,
22 tokenize=False,
23 add_generation_prompt=True
24)
25model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
26
27generated_ids = model.generate(
28 **model_inputs,
29 max_new_tokens=1024
30)
31generated_ids = [
32 output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
33]
34
35response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
36print(response)| Benchmark | Yi-1.5-9B-Chat | Mistral-Nemo-Instruct-2407 (12B) | Gemma-2-9b-it | Falcon3-10B-Instruct |
|---|---|---|---|---|
| IFEval | 60.46 | 63.80 | 74.36 | 78.17 |
| BBH (3-shot) | 36.95 | 29.68 | 42.14 | 44.82 |
| MATH Lvl-5 (4-shot) | 12.76 | 6.50 | 0.23 | 25.91 |
| GPQA (0-shot) | 11.30 | 5.37 | 14.77 | 10.51 |
| MUSR (0-shot) | 12.84 | 8.48 | 9.74 | 13.61 |
| MMLU-PRO (5-shot) | 33.06 | 27.97 | 31.95 | 38.10 |
| Category | Benchmark | Yi-1.5-9B-Chat | Mistral-Nemo-Instruct-2407 (12B) | Falcon3-10B-Instruct |
|---|---|---|---|---|
| General | MMLU (5-shot) | 68.8 | 66.0 | 73.9 |
| MMLU-PRO (5-shot) | 38.8 | 34.3 | 44 | |
| IFEval | 57.8 | 63.4 | 78 | |
| Math | GSM8K (5-shot) | 77.1 | 77.6 | 84.9 |
| GSM8K (8-shot, COT) | 76 | 80.4 | 84.6 | |
| MATH Lvl-5 (4-shot) | 3.3 | 5.9 | 22.1 | |
| Reasoning | Arc Challenge (25-shot) | 58.3 | 63.4 | 66.2 |
| GPQA (0-shot) | 35.6 | 33.2 | 33.5 | |
| GPQA (0-shot, COT) | 16 | 12.7 | 32.6 | |
| MUSR (0-shot) | 41.9 | 38.1 | 41.1 | |
| BBH (3-shot) | 50.6 | 47.5 | 58.4 | |
| CommonSense Understanding | PIQA (0-shot) | 76.4 | 78.2 | 78.4 |
| SciQ (0-shot) | 61.7 | 76.4 | 90.4 | |
| Winogrande (0-shot) | - | - | 71 | |
| OpenbookQA (0-shot) | 43.2 | 47.4 | 48.2 | |
| Instructions following | MT-Bench (avg) | 8.3 | 8.6 | 8.2 |
| Alpaca (WC) | 25.8 | 45.4 | 24.7 | |
| Tool use | BFCL AST (avg) | 48.4 | 74.2 | 90.5 |
| Code | EvalPlus (0-shot) (avg) | 69.4 | 58.9 | 74.7 |
| Multipl-E (0-shot) (avg) | - | 34.5 | 45.8 |
@misc{Falcon3,
title = {The Falcon 3 family of Open Models},
author = {TII Team},
month = {December},
year = {2024}
}| Metric | Value |
|---|---|
| Avg. | 35.19 |
| IFEval (0-Shot) | 78.17 |
| BBH (3-Shot) | 44.82 |
| MATH Lvl 5 (4-Shot) | 25.91 |
| GPQA (0-shot) | 10.51 |
| MuSR (0-shot) | 13.61 |
| MMLU-PRO (5-shot) | 38.10 |