The Italian dataset with the highest density of useful information per token. Built by ModotAI for training Italian language models.
Subset
File
Documents
Words
Description
Web Crawl
icc-web.parquet
~27K
~17M
Italian sources: news, tech, science, culture, law, food, sport
Wikipedia IT
wiki-it-clean.parquet
~1.35M
~698M
Cleaned Italian Wikipedia — removed Notes, Bibliography, Voci correlate, stub articles
Total: 1,377… See the full description on the dataset page:
https://huggingface.co/datasets/ThingAI/Italian-Common-Corpus.