FineWeb2-Ro-BERT is a large-scale pretraining dataset in the Romanian language. The data is derived from FineWeb2 and annotated using a bert architecture for signals such as educational quality or topic. More details can be found here.
Massive Scale: Contains approximately 54.1M rows (documents or sequences), providing comprehensive linguistic coverage for training robust Romanian embeddings and encoders.
You can load… See the full description on the dataset page:
https://huggingface.co/datasets/OpenLLM-Ro/fineweb2-ro-bert.