This checkpoint is from the clean pretraining line, stopped around the ~68k step range.
The data mix was:
TinyStories for compact story structure
filtered Cosmopedia-style story/explanation text
COCO captions for visual grounding
Flickr captions for additional image-caption variety
Current Behavior
The model is starting to learn some visual grounding. In simple manual tests, color words and broad scene cues started to affect generations, for example red images increasing red-related text.
The vision side is still pretty bad. It is weak at object identity, counting, spatial details, and robust visual question answering.
Use this as a toy research checkpoint, not a reliable assistant.