This repository contains the input scores dataset used for training M-FLYT as described in the paper Filter Like You Test: Data-Driven Data Filtering for CLIP Pretraining. The scores are formatted as a parquet dataset, and can be used to reproduce our results or to improve them by adding more or better scoring methods.
For code to use these scores and more information visit our GitHub repository.