Have dataset items that are somewhat evenly of each type. The LLM validation loss went down and then continue to rise afterwards,
so I guess the complexity of the dataset was too high.
The image sizes are between 1 and 10 pixels.
Here the majority of dataset items are histograms.
Smaller images. Here the image sizes are between 1 and 5 pixels.
Let's see if the LLM does better on this one.
Focus on pair comparisons finding color… See the full description on the dataset page:
https://huggingface.co/datasets/neoneye/simon-arc-task-v10.