This dataset is a multilingual subset of the P3 component of the original FLAN collection.
The original FLAN collection is extremely large and aggregates many instruction-following datasets across tasks and domains (see the original FLAN v2 repo for more information).
In contrast, this release contains a selected subset of the original FLAN data. The selection strategy mirrors the filtering and sampling used in the flan_v2_converted dataset.
The… See the full description on the dataset page:
https://huggingface.co/datasets/VillanovaAI/Multi-FLAN-P3.