Zagreus 0.4B xmoons improved
Zagreus 0.4B xmoons improved is an Italian multiple-choice model fine-tuned
from
mii-llm/zagreus-0.4B-ita.
It is the selected checkpoint from the reproducible H100 run that reached
44.35% accuracy on the official 10,000-question ITALIC benchmark in fast,
five-shot mode through vLLM.
The model was developed for the
mii-llm/Post-Training-Challenge.
The model is a compact Llama-style causal language model with 437,760,960
parameters. Its tokenizer and Llama 3 chat template come from
swap-uniba/LLaMAntino-3-ANITA-8B-Inst-DPO-ITA.
The repository is a self-contained Transformers checkpoint and uses
<|eot_id|> as both the end-of-turn and generation stop token.
Intended use
The model is intended for:
- research on compact Italian language models;
- Italian multiple-choice question answering;
- participation in the MII Post-Training Challenge;
- reproduction and analysis of the ITALIC post-training experiment.
It is not designed as a general-purpose assistant, a factual authority, or a
component for medical, legal, financial, or other high-stakes decisions.
Usage
Install PyTorch and Transformers, then load the checkpoint from the repository
root:
1import torch
2from transformers import AutoModelForCausalLM, AutoTokenizer
3
4model_id = "fgoose180/zagreus-0.4B-xmoons-improved"
5device = "cuda" if torch.cuda.is_available() else "cpu"
6dtype = torch.bfloat16 if device == "cuda" else torch.float32
7
8tokenizer = AutoTokenizer.from_pretrained(model_id)
9model = AutoModelForCausalLM.from_pretrained(
10 model_id,
11 dtype=dtype,
12).to(device)
13
14messages = [
15 {"role": "system", "content": "Sei un assistente utile."},
16 {
17 "role": "user",
18 "content": """Rispondi alla seguente domanda a scelta multipla sull'argomento 'data_geography'. La tua risposta deve essere nel seguente formato: 'LETTERA' (senza virgolette) dove LETTERA è una tra ABCD. Scrivi solo la lettera corrispondente alla tua risposta senza spiegazioni.
19
20Qual è la capitale d'Italia?
21
22A) Roma
23B) Milano
24C) Torino
25D) Napoli
26
27Risposta:""",
28 },
29]
30
31inputs = tokenizer.apply_chat_template(
32 messages,
33 add_generation_prompt=True,
34 return_tensors="pt",
35 return_dict=True,
36).to(device)
37
38with torch.inference_mode():
39 output = model.generate(
40 **inputs,
41 max_new_tokens=8,
42 do_sample=False,
43 eos_token_id=tokenizer.eos_token_id,
44 pad_token_id=tokenizer.pad_token_id,
45 )
46
47answer = tokenizer.decode(
48 output[0, inputs["input_ids"].shape[1] :],
49 skip_special_tokens=True,
50).strip()
51print(answer)
The reported benchmark result uses the five fixed official ITALIC
demonstrations. Zero-shot use, different prompts, sampling, or another chat
template should not be expected to reproduce the reported score.
Model architecture
| Property | Value |
|---|
| Architecture | LlamaForCausalLM |
| Model type | Llama-style decoder-only transformer |
| Parameters | 437,760,960 |
| Hidden size | 960 |
| Layers | 32 |
| Attention heads / KV heads | 15 / 5 |
| Vocabulary size | 128,256 |
| Maximum trained context | 2,048 tokens during SFT |
| Checkpoint serialization | SafeTensors |
Training
Data
The final training stream was built online from pinned revisions of four public
datasets:
| Dataset | Main contribution |
|---|
efederici/pinocchio | Italian language, culture, and general-knowledge MCQs |
FinancialSupport/italic_sft | ITALIC-style Italian MCQs |
FinancialSupport/italic_sft_ext | Extended ITALIC-style MCQs |
FinancialSupport/quiz_militare | Italian civic and general-knowledge MCQs |
Rows were normalized, exactly deduplicated, split before upsampling, and
decontaminated against the official ITALIC test and five demonstrations using
character TF-IDF cosine and word-shingle MinHash/LSH. The final artifacts were:
| Artifact | Rows | Unique questions |
|---|
| Clean training stream | 108,288 | 67,654 |
| Clean validation holdout | 6,745 | 6,745 |
The training stream contains 53,591 language draws and 54,697 culture draws.
The benchmark labels were not used for checkpoint selection.
Objective and hyperparameters
The model was trained for one epoch with completion-only supervised loss. Each
target was rendered with five distinct training-pool demonstrations and
independently permuted answer options. The official five ITALIC demonstrations
were reserved for validation and final evaluation.
| Parameter | Value |
|---|
| Hardware | 1x NVIDIA H100 |
| Optimizer steps | 6,768 |
| Batch size | 16 |
| Initial learning rate | 3e-4 |
| Schedule | 50-step warmup, cosine decay |
| Weight decay | 0.0 |
| Maximum sequence length | 2,048 |
| Forward precision | BF16 autocast |
| Parameters and optimizer updates | FP32 |
| Gradient clipping | 1.0 |
| Random seed | 0 |
| Few-shot probability / count | 1.0 / 5 |
| Option permutation probability | 1.0 |
The detached Modal pipeline took approximately 84 minutes end to end. Carbon
emissions were not measured, so no emissions estimate is reported.
Evaluation
The selected checkpoint was chosen at step 6,768 using only a clean validation
holdout and the criterion
0.5 * language_accuracy + 0.5 * culture_accuracy.
| Evaluation | Questions | Accuracy | Balanced accuracy | Unparsed |
|---|
| Validation holdout | 6,745 | 55.97% | 61.38% | 0 |
| Official ITALIC fast five-shot, vLLM 0.26 | 10,000 | 44.35% | not reported | 0 |
The official evaluation used greedy decoding, the official fast answer
extractor, a pinned ITALIC harness, and the official five demonstrations. The
result contained 4,435 correct answers out of 10,000.
Reproducibility
The training and evaluation implementation is available in
mattiacurri/zagreus-italic-challenge-xmoons.
The exact local source snapshot used for this release is identified by Git
commit
afc6b7f09a1e107f38ae04358ece7ed85f6be7a3.
Selected artifact identities:
1d679f3e79221aea95c64ef0be177530b02dce25d4745b0247cbae7b677f1ceab training pool
2749dbcdea8b244c3609acecddf618d9780559f501da18b09f650dfe10f2fb881 validation pool
30d534fd8eceec72b4fc3179d77afe7d6291766f5831bc4c8b69b9c45463c1656 selection.json
42ac2f6e412ea1e8f67ce6d8395c1c9845270101590d84de1be748e5919833c93 official result JSON
The source repository contains the full experiment and reproducibility reports.
Limitations and risks
- The model is optimized for Italian multiple-choice prompts and often emits
only an answer letter. It is not a broadly instruction-tuned chat model.
- The 44.35% figure is one benchmark result under one exact prompt and backend
configuration; it does not measure general Italian language competence.
- Exact and fuzzy decontamination reduce known overlap but cannot prove the
absence of semantic contamination or benchmark-distribution overfitting.
- The public training sources may contain factual errors, social biases,
stereotypes, or outdated information that can be inherited by the model.
- Outputs outside the trained answer format may be unreliable.
License and attribution
The base model weights are published under Apache-2.0. The included tokenizer
and chat template come from a Llama 3 derivative and remain subject to the
published Llama 3 terms; the metadata therefore uses the more restrictive
llama3 license identifier. Training datasets and benchmark assets retain
their own terms. Some FinancialSupport dataset cards do not declare a license,
so users should verify those terms before redistribution or commercial use.
Acknowledgements