This is a processed version of the Romanian Wikipedia subset from the FineWiki dataset, optimized for language model training and analysis. The dataset has been filtered to essential fields and enriched with token counts using tiktoken's cl100k_base encoding.
Total Pages: 493,462
Total Size: ~669 MB (compressed parquet)
Language: Romanian (ro)
Total Tokens: ~395 million tokens (cl100k_base encoding)
Source: FineWiki… See the full description on the dataset page:
https://huggingface.co/datasets/Yxanul/Romanian-Wiki.