This dataset contains domain names and counts of (non-deduplicated) URLs for every record in the CC-MAIN-2023-40 snapshot of the Common Crawl. It was collected from the AWS S3 version of Common Crawl via Amazon Athena.
This dataset is derived from Common Crawl data and is subject to Common Crawl's Terms of Use:
https://commoncrawl.org/terms-of-use.