This dataset is a metadata-enriched multilingual sample of the original FineWeb-2 dataset.
FineWeb-2 is a large-scale, high-quality web text corpus derived from Common Crawl, built through extensive filtering, deduplication, and language identification steps to support the training of modern large language models. It emphasizes text quality, diversity, and transparency, and includes rich crawl-level metadata inherited from Common Crawl.… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/fineweb-2-sample-60k-meta.