Villanova-2B-2603 is a fully open, multilingual instruction-tuned Large Language Model developed by Villanova.AI. Part of the Villanova project, it is designed to advance open European language technology with native support for five European languages. All model weights, training data sources, and training details are publicly released.
Built on top of Villanova-2B-Base-2603 — a 2.4B-parameter model pretrained from scratch — this instruction-tuned model offers strong multilingual instruction following and safety alignment under a fully open Apache 2.0 license.
Villanova-2B-2603 was extensively evaluated across 25 benchmarks covering Reasoning, Question Answering, Safety, and Instruction Following in both English and multilingual settings. All evaluations were performed using identical settings and prompts for fair comparison.
Tables are sorted by the main metric (descending). Models are grouped into Fully Open and Open Weight categories.
Overall Performance
Villanova-2B-2603 is the #1 fully open model in overall average across all benchmarks.
Model
Size
Reasoning
QA
Safety
Instr. Follow
Overall
Fully Open
Villanova-2B-2603
2.4B
31.0
33.1
39.5
45.1
36.9
OLMo-2-0425-1B-Instruct
1.2B
38.7
35.6
19.4
39.3
33.9
Minerva-7B-instruct-v1.0
7.4B
27.1
36.2
30.1
16.9
28.5
EuroLLM-1.7B-Instruct
1.7B
26.0
24.7
3.8
19.5
19.5
salamandra-2b-instruct
2.3B
23.6
26.6
9.6
15.7
20.0
Open Weight
Llama-3.2-3B-Instruct
3.2B
51.2
48.1
56.8
48.1
50.4
Qwen2.5-3B-Instruct
3.1B
39.4
35.8
54.7
46.8
42.9
Llama-3.2-1B-Instruct
1.2B
37.5
38.1
56.6
35.5
41.1
gemma-3-1b-it
1.0B
28.5
27.0
53.6
39.9
35.7
Qwen3-1.7B
1.7B
37.4
37.5
2.6
19.5
26.2
Instruction Following
Villanova-2B-2603 is the #1 fully open model for instruction following, and is competitive with larger open weight models. The MARCO benchmark evaluates structured instruction following across all five languages.
Model
Size
IFEval
MARCO-EN
MARCO-DE
MARCO-ES
MARCO-FR
MARCO-IT
Avg
Fully Open
Villanova-2B-2603
2.4B
62.0
39.4
40.5
44.2
42.5
42.1
45.1
OLMo-2-0425-1B-Instruct
1.2B
77.9
52.9
23.1
29.0
27.9
24.9
39.3
EuroLLM-1.7B-Instruct
1.7B
34.5
18.3
15.9
15.9
17.4
15.2
19.5
Minerva-7B-instruct-v1.0
7.4B
29.6
17.0
12.2
13.9
13.9
15.0
16.9
salamandra-2b-instruct
2.3B
26.4
17.7
12.2
12.0
12.9
12.9
15.7
Open Weight
Llama-3.2-3B-Instruct
3.2B
82.2
54.0
39.9
38.8
37.5
35.9
48.1
Qwen2.5-3B-Instruct
3.1B
71.5
47.3
37.5
42.5
41.0
40.7
46.8
gemma-3-1b-it
1.0B
74.5
42.7
27.5
33.3
27.9
33.3
39.9
Llama-3.2-1B-Instruct
1.2B
64.8
43.2
25.3
29.0
24.2
26.6
35.5
Qwen3-1.7B
1.7B
48.4
27.4
8.9
10.3
13.1
9.1
19.5
Key insight: While some models score higher on English-only IFEval, Villanova-2B-2603 delivers the most balanced multilingual instruction following, with MARCO scores of 40-44 across DE, ES, FR, IT. This is far ahead of OLMo (19-25) and Gemma (27-33) on non-English languages.
Safety (M-ALERT)
Villanova-2B-2603 is the #1 fully open model for safety. Safety was evaluated using the M-ALERT benchmark across all five languages.