Training the
Co-Instruct-562K dataset with LLaVA-1.5-7B to facilitate users that prefer the LLaVA structure.
We are working on improving it in the future but we also warn that this structure (direct projection) might not be very friendly to multi-image scenarios.