A reproducible dataset for studying how news-derived search terms propagate through multiple search engines, and how technical/content/accessibility signals vary across ranked results.
This repository contains:
the query/keyword seeds,
SERP indexing outputs,
the derived feature dataset used for analysis,
and documentation reports describing acceptance, outliers, and dataset health.
What’s inside (high level)
Core inputs… See the full description on the dataset page: https://huggingface.co/datasets/goker/comp-serp-data-v2.