This dataset is a pre-tokenized version of the Skylion007/openwebtext dataset
using the gemma tokenizer. As such, this dataset follows the same licensing as the original openwebtext dataset.
This pre-tokenization is done as a performance optimization for using the openwebtext dataset with a Gemma model (gemma-2b, gemma-2b-it, gemma-7b, gemma-7b-it).
This dataset was created using SAELens, with the following settings:… See the full description on the dataset page:
https://huggingface.co/datasets/Marlon154/openwebtext-gemma-2-context-128.