This dataset contains images and alt text from various sources.
It is used to control the quality of
https://huggingface.co/Mozilla/distilvit using the
https://github.com/mozilla/checkvite application
This application let users try out the model on the images and classify them. The dataset is then updated.
When an image is marked as need_training it will be use to fine-tune the model to fix some of its inaccuracies