Token counts for all datasets used in Marin pretraining runs.
web — Quality-classified Common Crawl text (Nemotron-CC)
code — Source code and… See the full description on the dataset page:
https://huggingface.co/datasets/marin-community/token-counts.