A curated list of multilingual text datasets available on Huggingface, designed to help users easily find datasets by language—including those for low-resource languages.
This index aims to make it easier to find datasets by language, addressing the common issue of inconsistent or unclear language codes across different datasets.
language: The English name… See the full description on the dataset page:
https://huggingface.co/datasets/agentlans/multilingual-dataset-index.