We introduce two augmented long-context multimodal datasets: Long Visual Question Answering and Long multimodal Retrieval. These datasets aim to enhance VLMs' long-context training and establish a systematic evaluation framework, thereby addressing the challenges associated with long-context understanding that extend beyond the scope of existing training data.