A large-scale, sector-labeled corpus of websites designed for web-tailored multi-class classification.
SoAC (Sector of Activity Corpus) comprises 195,495 unique websites collected in 2024 and annotated with both coarse- and fine-grained sector labels. It was created to benchmark sector-based website classification and to support research in privacy policy analysis, regulatory compliance, and targeted content… See the full description on the dataset page:
https://huggingface.co/datasets/Shahriar/SoAC_Corpus.