This dataset comes from HuggingFaceFW/fineweb-2 and HuggingFaceFW/fineweb-edu. It includes five languages: Vietnamese, English, French, Japanese, and Chinese. Each language has 200,000 training samples and 10,000 test samples, totaling 1 million rows for training and 50,000 rows for testing.
The default subset is the combination of other subset