The π· FineWeb dataset consists of more than 15T tokens of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the π datatrove library, our large scale data processing library.
π· FineWeb was originally meant to be a fully open replication of π¦
RefinedWeb, with a release of the full dataset under⦠See the full description on the dataset page:
https://huggingface.co/datasets/akhilhsingh/homeo-dataset.