Views
No views yet
nvidia/Cosmos3-Nano omnimodal
world model, created with bitsandbytes.
Only the large Cosmos3OmniTransformer is quantized; the VAE and the
text/sound tokenizers are bundled unchanged at bf16, so the repo is
self-contained and fully drop-in.| Property | Value |
|---|---|
| Repo size | 11 GB (vs ~34 GB bf16) |
| Quantized component | transformer — 8.3 GB NF4 (vs ~32 GB bf16) |
| Quantization | NF4 (bitsandbytes), double quantization, bnb_4bit_compute_dtype=bfloat16 |
| Modes | text-to-image, text-to-video, image-to-video (+ optional sound) |
| Base params | 16B (omnimodal) |
| VRAM (loaded) | ~11 GB |
| Source weights | nvidia/Cosmos3-Nano (bf16) |
| Tested on | NVIDIA GB10 (DGX Spark) |
diffusers build with Cosmos 3 support (currently from source) plus
bitsandbytes. The NF4 config is embedded — do not pass a
quantization_config, and do not call .to(dtype) on a 4-bit model.pip install "git+https://github.com/huggingface/diffusers.git" bitsandbytes accelerate1import torch
2from diffusers import Cosmos3OmniPipeline
3
4pipe = Cosmos3OmniPipeline.from_pretrained(
5 "SanDiegoDude/Cosmos3-Nano-nf4",
6 torch_dtype=torch.bfloat16,
7 enable_safety_checker=False, # skips the optional cosmos_guardrail dependency
8).to("cuda")
9
10result = pipe("A small warehouse robot beside a blue box, clean studio lighting.")
11frames = result.video[0] # text-to-image returns a single frame
12frames[0].save("out.png")