A uniformly randomized subset of HuggingFaceFW/finewiki, created to provide a smaller and more manageable dataset for analysis, fine-tuning, and benchmarking.
This sample includes Wikipedia articles from languages with more than one million pages. Sampling is performed uniformly at random instead of alphabetically to ensure unbiased representation.
Languages were selected based on page count and… See the full description on the dataset page:
https://huggingface.co/datasets/agentlans/HuggingFaceFW-finewiki-sample.