This dataset is a pre-tokenized version of the Skylion007/openwebtext dataset
using the llama3 tokenizer. As such, this dataset follows the same licensing as the original openwebtext dataset.
This pre-tokenization is done as a performance optimization for using the openwebtext dataset with a Llama3 model.
This dataset was created using SAELens, with the following settings:
context_size: 8192
shuffled: true
begin_batch_token: "bos"… See the full description on the dataset page:
https://huggingface.co/datasets/chanind/openwebtext-llama3.