A human-curated benchmark dataset for evaluating web scraping engines on content quality.
This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time.
id: Sequential identifier
url:… See the full description on the dataset page:
https://huggingface.co/datasets/firecrawl/scrape-content-dataset-v1.