A collection of all books written in Italian that appear on Project Gutenberg, meant for pretraining of Large Language Models.
We collected the plaintext version of each book and lightly processed it to remove licensing text that is usually included in the original text of Project Gutenberg.
Simply download the dataset as you would with any other Hugging Face dataset:
from datasets import load_dataset
dataset =… See the full description on the dataset page:
https://huggingface.co/datasets/tommasobonomo/ITAGutenberg.