These are the training datasets for the Goldfish models, as described in our paper, Goldfish: Monolingual Language Models for 350 Languages (Chang et al., 2026).
Along with citing the Goldfish paper, if using this dataset, we encourage researchers to cite the individual datasets listed in our paper.
@inproceedings{chang-etal-2026-goldfish,
title={Goldfish: Monolingual Language Models for 350 Languages},
author={Chang, Tyler A. and Arnett… See the full description on the dataset page:
https://huggingface.co/datasets/goldfish-models/fish-food.