In this repository, we specifically provide the 5k training samples from the complete GameQA-140K dataset used in our work for GRPO training of the models.
Refer to our paper for details. And our code for training and evaluation is at
https://github.com/tongjingqi/Code2Logic.
This is the first work, to the best of our knowledge, that leverages game code to synthesize multimodal reasoning data for… See the full description on the dataset page:
https://huggingface.co/datasets/OpenMOSS-Team/GameQA-5K.