Tokenized sentence/paragraph segments of OpenWebText (Skylion007/openwebtext, revision main), encoded with GPT-2 BPE plus three extra special tokens, then assigned to length buckets.
This repository is a derived dataset. It does not rediscover or replace the original text corpus. Every input_ids sequence comes from documents in OpenWebText.
Packaging / conversion code & this card: MIT (see LICENSE)
Underlying web… See the full description on the dataset page:
https://huggingface.co/datasets/fengluoqiuwu/owt-bucket.