Work in progress. Coverage is partial — additional CommonCrawl dumps are being embedded and uploaded incrementally.
Pre-computed dense and sparse embeddings for the FineWeb web corpus, ready for direct ingestion into a vector database (e.g. Qdrant).
FineWeb is a 15-trillion-token English web dataset derived from 96 CommonCrawl snapshots spanning Summer 2013 through June 2025. It was produced by Hugging… See the full description on the dataset page:
https://huggingface.co/datasets/nleroy917/fineweb-bge-large-en-v1.5.