π Website | π₯οΈ Code | π Paper
A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 1742 billion tokens of science, technology, engineering, and mathematics content.
This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that require⦠See the full description on the dataset page:
https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm.