The dataset is built from the redpajamas dataset after filtering by marketing keywords list that can be found here
The full scripts to recreate the raw dataset before sharding can be found here.
The dataset includes:
~4.8B tokens from raw contents.
To start exploring and get to know the dataset you can run the script:
import datasets
for sample in ds:… See the full description on the dataset page:
https://huggingface.co/datasets/marketeam/raw_redpajamas.