10K lines of Latin Script languages from FineWeb2 for lightweight tokenizer adaptation experiments.
ds_ibo = load_dataset(
"MultilingualUnigramLM/FineWeb2-10K",
split="ibo_Latn"… See the full description on the dataset page:
https://huggingface.co/datasets/MultilingualUnigramLM/FineWeb2-10K.