Packed pretraining corpus, 2,050,000 rows x 4096 tokens = 8.397B tokens, tokenized with
fhai50032/QTK-81K.
There is no attention_mask column: the corpus is packed, so every position is a real token and
the mask would be all ones on every row.… See the full description on the dataset page:
https://huggingface.co/datasets/fhai50032/BiBo_corpus-4096-packed-qtk-2.05M-hi35-en65.