This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs.
Source
Samples
Description… See the full description on the dataset page:
https://huggingface.co/datasets/Amshaker/Mobile-O-Pre-Train.