This repository contains the llava-cc3m-smolRGPT dataset, a key component of the research presented in the paper SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters.
Code Repository:
https://github.com/abtraore/SmolRGPT
Recent advances in vision-language models (VLMs) have enabled powerful multimodal reasoning, but state-of-the-art approaches typically rely on extremely large models with… See the full description on the dataset page:
https://huggingface.co/datasets/Abdrah/llava-cc3m-smolRGPT.