Manu-FineWeb is a high-quality, large-scale corpus specifically curated for the manufacturing domain. It was extracted from the 15-trillion-token FineWeb dataset and refined to facilitate efficient domain-specific pretraining for models like ManufactuBERT.
Developed by: Robin Armingaud and Romaric Besançon (Université Paris-Saclay, CEA, List)
Statistics: 2B tokens/4,5 million documents
The dataset was… See the full description on the dataset page:
https://huggingface.co/datasets/rarmingaud/Manu-FineWeb.