A multilingual adaptation of SciRIFF extending a filtered subset of the original English-only instruction-following scientific literature dataset to five languages with permissively licensed synthetic translations.
The original SciRIFF dataset, by AllenAI, includes ~137 K instruction-following demonstrations for 54 scientific literature understanding tasks, organized with rich metadata describing domains, task families, and context. It was developed as a benchmark for… See the full description on the dataset page:
https://huggingface.co/datasets/VillanovaAI/Multi-SciRIFF.