A T2I Adapter trained on VJEPA 2.1 embeddings, to reconstruct the embeddings back into pixel space.
Some Examples:
P.S. Most of these examples were generated using the LCM Scheduler + LCM LoRA for 2-4 step per frame generations. Using DDIM would yield much better results, and generate much more faithful colors
Even facial details are conserved if the face occupies a significant part of the frame