This dataset is a collection of two datasets provided by Marc A. Kastner on GitHub.
I merged the datasets and kept only the word, visual, phonetic, and textual columns.
The data is scaled using a MinMaxScaler so that the whole dataset can be used as one.
This dataset is ideal for training and evaluating machine learning models for word imageability.
We extend our heartfelt gratitude to all the authors of the… See the full description on the dataset page:
https://huggingface.co/datasets/StephanAkkerman/imageability-corpus.