PRX (Photoroom Experimental) is a 1.3-billion-parameter text-to-image model trained entirely from scratch and released under an Apache 2.0 license.
It is part of Photoroom’s broader effort to open-source the complete process behind training large-scale text-to-image models — covering architecture design, optimization strategies, and post-training alignment. The goal is to make PRX both a strong open baseline and a transparent research reference for those developing or studying diffusion-transformer models.
PRX is designed to be lightweight yet capable, easy to fine-tune or extend, and fully open.
PRX generates high-quality images from text using a simplified MMDiT architecture where text tokens don’t update through transformer blocks. It uses flow matching with discrete scheduling for efficient sampling and Google’s T5-Gemma-2B-2B-UL2 model for multilingual text encoding. The model has around 1.3B parameters and delivers fast inference without sacrificing quality. You can choose between Flux VAE for balanced quality and speed, or DC-AE for higher latent compression and faster processing.
This card in particular describes Photoroom/prx-256-t2i, one of the PRX model variants:
1from diffusers.pipelines.prx import PRXPipeline
23pipe = PRXPipeline.from_pretrained(4"Photoroom/prx-256-t2i",5 torch_dtype=torch.bfloat16
6).to("cuda")78prompt ="A front-facing portrait of a lion in the golden savanna at sunset"9image = pipe(prompt, num_inference_steps=28, guidance_scale=5.0).images[0]10image.save("lion.png")
PRX models were trained from scratch using recent advances in diffusion and flow-matching training. We experimented with a range of modern techniques for efficiency, stability, and alignment, which we’ll cover in more detail in our upcoming series of research posts: