This is the vocabulary of the 10Bt sample of FineWeb. This vocabulary was obtained by normalizing and pretokenizing the vocabulary using the bert-base-uncased tokenizer.
You can use this vocabulary to:
Obtain probabilities of subparts of your corpus.
Define useful tokenizer extensions without fitting a new tokenizer.
Analyzing the semantic content of the corpus
The dataset consists of 1.85 million tokens with their associated frequency and… See the full description on the dataset page:
https://huggingface.co/datasets/stephantulkens/fineweb-vocab.