This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages.
The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page:
https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.