The dataset-eval dataset is a multilingual and multi-domain dataset designed for evaluating language model performance during training. It can be used for
performance tracking, generalization diagnostics across languages or domains, and for implementing early stopping mechanisms.
The examples included were automatically selected as High quality by the EuroBERT-210m-Quality model,
trained to estimate web text quality in multiple… See the full description on the dataset page: https://huggingface.co/datasets/TempestTeam/dataset-eval.