This repository hosts BERTino, an Italian DistilBERT model pre-trained by
indigo.ai
on a large general-domain Italian corpus. BERTino is task-agnostic and can be
fine-tuned for every downstream task.
The pre-training corpus that we used is the union of the
Paisa and
ItWaC
corpora. The final corpus counts 14 millions of sentences for a total of 12 GB
of text.
To validate the pre-training that we conducted, we evaluated BERTino on the
Italian ParTUT,
Italian ISDT,
Italian WikiNER
and multi-class sentence classification tasks. We report for comparison results
obtained by the
teacher model
fine-tuned in the same tasks and for the same number of epochs.