2024-10-28: We release new version of inference code, optimizing the memory usage and time cost. You can refer to docs/inference.md for detailed information.
2024-10-22: :fire: We release the first version of OmniGen. Model Weight: Shitao/OmniGen-v1 HF Demo: 🤗
2. Overview
OmniGen is a unified image generation model that can generate a wide range of images from multi-modal prompts. It is designed to be simple, flexible, and easy to use. We provide inference code so that everyone can explore more functionalities of OmniGen.
Existing image generation models often require loading several additional network modules (such as ControlNet, IP-Adapter, Reference-Net, etc.) and performing extra preprocessing steps (e.g., face detection, pose estimation, cropping, etc.) to generate a satisfactory image. However, we believe that the future image generation paradigm should be more simple and flexible, that is, generating various images directly through arbitrarily multi-modal instructions without the need for additional plugins and operations, similar to how GPT works in language generation.
Due to the limited resources, OmniGen still has room for improvement. We will continue to optimize it, and hope it inspires more universal image-generation models. You can also easily fine-tune OmniGen without worrying about designing networks for specific tasks; you just need to prepare the corresponding data, and then run the script. Imagination is no longer limited; everyone can construct any image-generation task, and perhaps we can achieve very interesting, wonderful, and creative things.
OmniGen is a unified image generation model that you can use to perform various tasks, including but not limited to text-to-image generation, subject-driven generation, Identity-Preserving Generation, image editing, and image-conditioned generation. OmniGen doesn't need additional plugins or operations, it can automatically identify the features (e.g., required object, human pose, depth mapping) in input images according to the text prompt.
We showcase some examples in inference.ipynb. And in inference_demo.ipynb, we show an interesting pipeline to generate and modify an image.
You can control the image generation flexibly via OmniGen
demo
If you are not entirely satisfied with certain functionalities or wish to add new capabilities, you can try fine-tuning OmniGen.
You also can create a new environment to avoid conflicts:
# Create a python 3.10.12 conda env (you could also use virtualenv)
conda create -n omnigen python=3.10.12
conda activate omnigen
# Install pytorch with your CUDA version, e.g.
pip install torch==2.3.1+cu118 torchvision --extra-index-url https://download.pytorch.org/whl/cu118
git clone https://github.com/staoxiao/OmniGen.git
cd OmniGen
pip install -e .
Here are some examples:
python
1from OmniGen import OmniGenPipeline
23pipe = OmniGenPipeline.from_pretrained("Shitao/OmniGen-v1")4# Note: Your local model path is also acceptable, such as 'pipe = OmniGenPipeline.from_pretrained(your_local_model_path)', where all files in your_local_model_path should be organized as https://huggingface.co/Shitao/OmniGen-v1/tree/main567## Text to Image8images = pipe(9 prompt="A curly-haired man in a red shirt is drinking tea.",10 height=1024,11 width=1024,12 guidance_scale=2.5,13 seed=0,14)15images[0].save("example_t2i.png")# save output PIL Image1617## Multi-modal to Image18# In the prompt, we use the placeholder to represent the image. The image placeholder should be in the format of <img><|image_*|></img>19# You can add multiple images in the input_images. Please ensure that each image has its placeholder. For example, for the list input_images [img1_path, img2_path], the prompt needs to have two placeholders: <img><|image_1|></img>, <img><|image_2|></img>.20images = pipe(21 prompt="A man in a black shirt is reading a book. The man is the right man in <img><|image_1|></img>.",22 input_images=["./imgs/test_cases/two_man.jpg"],23 height=1024,24 width=1024,25 guidance_scale=2.5,26 img_guidance_scale=1.6,27 seed=028)29images[0].save("example_ti2i.png")# save output PIL image
If out of memory, you can set offload_model=True. If the inference time is too long when inputting multiple images, you can reduce the max_input_image_size. For the required resources and the method to run OmniGen efficiently, please refer to docs/inference.md#requiremented-resources.
If you find this repository useful, please consider giving a star ⭐ and citation
@article{xiao2024omnigen,
title={Omnigen: Unified image generation},
author={Xiao, Shitao and Wang, Yueze and Zhou, Junjie and Yuan, Huaying and Xing, Xingrun and Yan, Ruiran and Wang, Shuting and Huang, Tiejun and Liu, Zheng},
journal={arXiv preprint arXiv:2409.11340},
year={2024}
}