Model type:
D-RT2-Style is one of the baselines in our LLaRA paper, following the style of
RT2.
This is an open-source visuomotor policy trained by fine-tuning
LLaVA-7b-v1.5 on instruction-following data
D-RT2-Style, converted from
VIMA-Data.
For the conversion code, please refer to
convert_vima.ipynb