๐ Search this corpus online: query it with sub-second full-text search and n-gram counts โ in the browser or via a public, keyless REST API, no download required โ at infini-news.uni-graz.at (API reference).
A multilingual news corpus extracted from
Common Crawl CC-News WARC files.
One row per article, with body text extracted via
trafilatura,
WARC provenance, and derived metadata (publish date, language, topic,
byte hashes) in a single flat schema. Coversโฆ See the full description on the dataset page:
https://huggingface.co/datasets/ruggsea/infini-news-corpus.