Lai Wei, Zihao Jiang, Weiran Huang, Lichao Sun.
Shanghai Jiao Tong University, Lehigh University
Paper, Link, Code
Multimodal large language models acquire their instruction-following capabilities through a two-stage training process: pre-training on image-text pairs and fine-tuning on supervised vision-language instruction data. Recent studies have shown that large language models can… See the full description on the dataset page:
https://huggingface.co/datasets/WaltonFuture/InstructionGPT-4.