The PixelByte model generates mixed sequences of text and images, handling transitions with line breaks and maintaining image dimension consistency.
We use the
PixelBytes-Pokemon dataset, available on Hugging Face:
PixelBytes-Pokemon. It contains text and image sequences of Pokémon for training our model.
Furfaro, F. (2024). PixelBytes: A Unified Multimodal Representation Learning Project. (
https://github.com/fabienfrfr/PixelBytes)