This is a dataset scraped from multiple news and entertainment websites.Each entry is labeled as Slop or Non-Slop depending on content quality.
š Dataset Details
Website: Source domain (e.g., NY Times, BBC)
Title: Title of the page
URL: Direct link
Domain: News, Lifestyle, etc.
Slop: Label (Slop / Non-Slop)
Content: Cleaned text from the HTML (for the version with the raw html check stop-slop-data-html)