The curation was done using qualitative analysis of the dataset, following visualization techniques like PCA and UMAP and score-based categorization of the samples using metrics like hardness, mistakenness, or uniqueness.
The code of the curation can be found on GitHub:👉
https://github.com/Conscht/MNIST_Curation_Repo/tree/main
This curated version of MNIST introduces an additional IDK (“I Don’t Know”) label for digits that are ambiguous, noisy… See the full description on the dataset page:
https://huggingface.co/datasets/Consscht/MNIST-Curation.