Views
No views yet

1from transformers import AutoTokenizer, AutoModelForCausalLM
2
3
4from transformers import AutoModelForCausalLM, AutoTokenizer
5
6model_name = "tiiuae/Falcon3-10B-Instruct"
7
8model = AutoModelForCausalLM.from_pretrained(
9 model_name,
10 torch_dtype="auto",
11 device_map="auto"
12)
13tokenizer = AutoTokenizer.from_pretrained(model_name)
14
15prompt = "How many hours in one day?"
16messages = [
17 {"role": "system", "content": "You are a helpful friendly assistant Falcon3 from TII, try to follow instructions as much as possible."},
18 {"role": "user", "content": prompt}
19]
20text = tokenizer.apply_chat_template(
21 messages,
22 tokenize=False,
23 add_generation_prompt=True
24)
25model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
26
27generated_ids = model.generate(
28 **model_inputs,
29 max_new_tokens=1024
30)
31generated_ids = [
32 output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
33]
34
35response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
36print(response)| Category | Benchmark | Yi-1.5-9B-Chat | Mistral-Nemo-Base-2407 (12B) | Falcon3-10B-Instruct |
|---|---|---|---|---|
| General | MMLU (5-shot) | 70 | 65.9 | 71.6 |
| MMLU-PRO (5-shot) | 39.6 | 32.7 | 44 | |
| IFEval | 57.6 | 63.4 | 78 | |
| Math | GSM8K (5-shot) | 76.6 | 73.8 | 83.1 |
| GSM8K (8-shot, COT) | 78.5 | 73.6 | 81.3 | |
| MATH Lvl-5 (4-shot) | 8.8 | 0.4 | 22.1 | |
| Reasoning | Arc Challenge (25-shot) | 51.9 | 61.6 | 64.5 |
| GPQA (0-shot) | 35.4 | 33.2 | 33.5 | |
| GPQA (0-shot, COT) | 16 | 12.7 | 32.6 | |
| MUSR (0-shot) | 41.9 | 38.1 | 41.1 | |
| BBH (3-shot) | 49.2 | 43.6 | 58.4 | |
| CommonSense Understanding | PIQA (0-shot) | 76.4 | 78.2 | 78.4 |
| SciQ (0-shot) | 61.7 | 76.4 | 90.4 | |
| Winogrande (0-shot) | - | - | 71.3 | |
| OpenbookQA (0-shot) | 43.2 | 47.4 | 48.2 | |
| Instructions following | MT-Bench (avg) | 8.28 | 8.6 | 8.17 |
| Alpaca (WC) | 25.81 | 45.44 | 24.7 | |
| Tool use | BFCL AST (avg) | 48.4 | 74.2 | 86.3 |
| Code | EvalPlus (0-shot) (avg) | 69.4 | 58.9 | 74.7 |
| Multipl-E (0-shot) (avg) | - | 34.5 | 45.8 |
@misc{Falcon3,
title = {The Falcon 3 family of Open Models},
author = {TII Team},
month = {December},
year = {2024}
}| Metric | Value |
|---|---|
| Avg. | 35.19 |
| IFEval (0-Shot) | 78.17 |
| BBH (3-Shot) | 44.82 |
| MATH Lvl 5 (4-Shot) | 25.91 |
| GPQA (0-shot) | 10.51 |
| MuSR (0-shot) | 13.61 |
| MMLU-PRO (5-shot) | 38.10 |