The dataset URLs and Domain Names are collected from the following sources:
mC4
Description: The Multilingual Colossal Common Crawl Corpus (mC4) is a cleaned version of the Common Crawl's web corpus, curated by the Allen Institute for Artificial Intelligence. It contains approximately 170 million URLs.
Source: mC4 Dataset on Hugging Face