pony-diffusion is a latent text-to-image diffusion model that has been conditioned on high-quality pony SFW-ish images through fine-tuning.
With special thanks to
Waifu-Diffusion for providing finetuning expertise and
Novel AI for providing necessary compute.
The model originally used for fine-tuning is an early finetuned checkpoint of
waifu-diffusion on top of
Stable Diffusion V1-4, which is a latent image diffusion model trained on
LAION2B-en.
This particular checkpoint has been fine-tuned with a learning rate of 5.0e-6 for 4 epochs on approximately 80k pony text-image pairs (using tags from derpibooru) which all have score greater than 500 and belong to categories safe or suggestive.
This model is open access and available to all, with a CreativeML OpenRAIL-M license further specifying rights and usage.
The CreativeML OpenRAIL License specifies:
This model can be used for entertainment purposes and as a generative art assistant.
1import torch
2from torch import autocast
3from diffusers import StableDiffusionPipeline, DDIMScheduler
4model_id = "AstraliteHeart/pony-diffusion"
5device = "cuda"
6pipe = StableDiffusionPipeline.from_pretrained(
7 model_id,
8 torch_dtype=torch.float16,
9 revision="fp16",
10 scheduler=DDIMScheduler(
11 beta_start=0.00085,
12 beta_end=0.012,
13 beta_schedule="scaled_linear",
14 clip_sample=False,
15 set_alpha_to_one=False,
16 ),
17)
18pipe = pipe.to(device)
19prompt = "pinkie pie anthro portrait wedding dress veil intricate highly detailed digital painting artstation concept art smooth sharp focus illustration Unreal Engine 5 8K"
20with autocast("cuda"):
21 image = pipe(prompt, guidance_scale=7.5)["sample"][0]
22
23image.save("cute_poner.png")
This project would not have been possible without the incredible work by the
CompVis Researchers.
In order to reach us, you can join our
Discord server.