GitHub:
https://github.com/commoncrawl/cc-host-index
Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The
information is aggregated from the Common Crawl columnar index,
web graph, and raw crawler logs.
DuckDB can read directly from Huggingface:
$ duckdb
D FROM… See the full description on the dataset page:
https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.