Dataset Card for OpenWebText n-grams
Dataset Summary
This dataset contains 246K of the most common token-based (GPT-2/GPT-3) n-grams (n=1 to n=6), in the OpenWebText (OWT) dataset.
For convenient searching, it provides full tokens/strings, as well as per-position tokens/strings.
Usage
Generally, this dataset allows identifying the most common n-grams in a text corpus.
When researching LLMs tokenized similarly to GPT-2/GPT-3, it allows: