CogAlign is a post-training strategy for Vision Language Models (VLMs) aimed at enhancing their visual arithmetic capabilities. This repository presents the training data for CogAlign, a synthetic dataset containing 64,000 examples designed to facilitate this post-training process.
CogAlign is inspired by Piaget's theory of cognitive development and focuses on improving a VLM's understanding of… See the full description on the dataset page:
https://huggingface.co/datasets/Salesforce/CogAlign.