1Stony Brook University 2University of Wisconsin-Madison
Model details
Model type:
LLaRA is an open-source visuomotor policy trained by fine-tuning LLaVA-7b-v1.5 on instruction-following data D-inBC and 4 auxiliary datasets, converted from VIMA-Data.
For the conversion code, please refer to convert_vima.ipynb
Model date:
llava-1.5-7b-llara-D-inBC-Aux-B-VIMA-80k was trained in June 2024.
Primary intended uses:
The primary use of LLaRA is research on large multimodal models for robotics.
Primary intended users:
The primary intended users of the model are researchers and hobbyists in robotics, computer vision, natural language processing, machine learning, and artificial intelligence.