This model is the first published base-pretraining artifact for the Sophira project.
Intended uses:
Italian language modeling research
downstream evaluation and benchmarking
initialization for later instruction tuning or task adaptation
reproducibility work around open Italian foundation-model pretraining
Out-of-Scope Use
This artifact is not yet documented or evaluated as suitable for:
safety-critical production deployment
legal, medical, or financial decision support
factual-reliability-sensitive assistant use without downstream evaluation
multilingual production use outside Italian-first evaluation
Training Data
Canonical training sources:
uonlp/CulturaX Italian subset
PleIAs/Italian-PD
Full-pass mixture used for this run:
82.977% CulturaX Italian
17.023% Italian-PD
Measured source-token counts:
CulturaX_it: 135.197.223.323
Italian_PD: 27.736.341.018
Total: 162.933.564.341
The project license policy and source-license references remain tracked in DATA_LICENSES.md.
Training Procedure
Training Framework: Megatron-LM
Runtime: .venv-apex
Cluster: CINECA Leonardo. We acknowledge the CINECA award under the ISCRA initiative, for the availability of high-performance computing resources and support.
Topology: 3 nodes / 12 GPUs
Sequence length: 2048
Micro-batch size: 3
Global batch size: 72
Tokens per step: 147.456
Target steps: 1.104.964
Checkpoint interval: 50.000
Final successful completion occurred on Sunday, July 26, 2026.
Evaluation
End-of-training validation:
iteration: 1.104.964
validation loss: 2.910531E+00
validation perplexity: 1.836655E+01
This is the validated Megatron end-of-training metric for the completed full-pass run.
Release validation and benchmark highlights:
Hugging Face controlled-generation summary:
empty_generation_rate: 0.0
avg_repeated_bigram_fraction: 0.1148
avg_repeated_trigram_fraction: 0.0809
avg_topic_keyword_overlap: 0.23
stronger benchmark groups (BLiMP-IT):
verbal_class_and_argument_structure: 0.9500
pronouns: 0.7188
agreement_and_inflection: 0.7454 with appended EOS