81,000 domains from 20 large website categories.
For each domain we collected a screenshot (1024x1024), the final document state (as HTML) and a PCAP file containing all the network traffic.
All of this occurred within a docker container so the network traffic is not cross-contaminated.
You can duplicate our work using our scraping tool) and the domain list in this repository.
Unfortunately, the domain categories have not been approved for release at this time
This… See the full description on the dataset page:
https://huggingface.co/datasets/ryanray-umich/Flint2025.