Views
No views yet

Abstract: Diffusion models have demonstrated promising performance in real-world video super-resolution (VSR). However, the dozens of sampling steps they require, make inference extremely slow. Sampling acceleration techniques, particularly single-step, provide a potential solution. Nonetheless, achieving one step in VSR remains challenging, due to the high training overhead on video data and stringent fidelity demands. To tackle the above issues, we propose DOVE, an efficient one-step diffusion model for real-world VSR. DOVE is obtained by fine-tuning a pretrained video diffusion model (i.e., CogVideoX). To effectively train DOVE, we introduce the latent–pixel training strategy. The strategy employs a two-stage scheme to gradually adapt the model to the video super-resolution task. Meanwhile, we design a video processing pipeline to construct a high-quality dataset tailored for VSR, termed HQ-VSR. Fine-tuning on this dataset further enhances the restoration capability of DOVE. Extensive experiments show that DOVE exhibits comparable or superior performance to multi-step diffusion-based VSR methods. It also offers outstanding inference efficiency, achieving up to a 28× speed-up over existing methods such as MGLD-VSR.



1# Clone the github repo and go to the default directory 'DOVE'.
2git clone https://github.com/zhengchen1999/DOVE.git
3conda create -n DOVE python=3.11
4conda activate DOVE
5pip install -r requirements.txt
6pip install diffusers["torch"] transformers
7pip install pyiqadatasets/train/.| Dataset | Type | # Videos / Images | Download |
|---|---|---|---|
| HQ-VSR | Video | 2,055 | Google Drive |
| DIV2K-HR | Image | 800 | Official Link |
1datasets/
2└── train/
3 ├── HQ-VSR/
4 └── DIV2K_train_HR/💡 HQ-VSR description:
- Construct using our four-stage video processing pipeline.
- Extract 2,055 videos from OpenVid-1M, suitable for video super-resolution (VSR) training.
- Detailed configuration and statistics are provided in the paper.
| Dataset | Type | # Num | Download |
|---|---|---|---|
| UDM10 | Synthetic | 10 | Google Drive |
| SPMCS | Synthetic | 30 | Google Drive |
| YouHQ40 | Synthetic | 40 | Google Drive |
| RealVSR | Real-world | 50 | Google Drive |
| MVSR4x | Real-world | 15 | Google Drive |
| VideoLQ | Real-world | 50 | Google Drive |
datasets/test/) is correct before running inference.1datasets/
2└── test/
3 └── [DatasetName]/
4 ├── GT/ # Ground Truth: folder of high-quality frames (one per clip)
5 ├── GT-Video/ # Ground Truth (video version): lossless MKV format
6 ├── LQ/ # Low-quality Input: folder of degraded frames (one per clip)
7 └── LQ-Video/ # Low-Quality Input (video version): lossless MKV format| Model Name | Description | HuggingFace | Google Drive | Baidu Disk | Visual Results |
|---|---|---|---|---|---|
| DOVE | Base version, built on CogVideoX1.5-5B; | TODO | Download | Download | Download |
| DOVE-2B | Smaller version, based on CogVideoX-2B | TODO | TODO | TODO | TODO |
Place downloaded model files into thepretrained_models/folder, e.g.,pretrained_models/DOVE.
Note: Training requires 4×A100 GPUs (80 GB each). You can optionally reduce the number of GPUs and use LoRA fine-tuning to reduce GPU memory requirements.
| Type | Dataset / Model | Path |
|---|---|---|
| Training | HQ-VSR, DIV2K-HR | datasets/train/ |
| Testing | UDM10 | datasets/test/ |
| Pretrained model | CogVideoX1.5-5B | pretrained_models/ |
1# 🔹 Train dataset
2python finetune/scripts/prepare_dataset.py --dir /data2/chenzheng/DOVE/datasets/train/HQ-VSR
3python finetune/scripts/prepare_dataset.py --dir /data2/chenzheng/DOVE/datasets/train/DIV2K_train_HR
4# 🔹 Testing dataset
5python finetune/scripts/prepare_dataset.py --dir /data2/chenzheng/DOVE/datasets/test/UDM10/GT-Video
6python finetune/scripts/prepare_dataset.py --dir /data2/chenzheng/DOVE/datasets/test/UDM10/LQ-Videofinetune/ directory and perform the first-stage training (latent-space) using:bash train_ddp_one_s1.shpython finetune/scripts/prepare_sft_ckpt.py --checkpoint_dir checkpoint/DOVE-s1/checkpoint-10000bash train_ddp_one_s2.shpython finetune/scripts/prepare_sft_ckpt.py --checkpoint_dir checkpoint/DOVE-/checkpoint-500💡 Prompt Optimization: DOVE uses an empty prompt (""). To accelerate inference, we pre-load the empty prompt embedding frompretrained_models/prompt_embeddings. When the prompt is empty, the pre-loaded embedding is used directly, bypassing text encoding and reducing overhead.
1# 🔹 Demo inference
2python inference_script.py \
3 --input_dir datasets/demo \
4 --model_path pretrained_models/DOVE \
5 --output_path results/DOVE/demo \
6 --is_vae_st \
7 --save_format yuv420p
8
9# 🔹 Reproduce paper results
10python inference_script.py \
11 --input_dir datasets/test/UDM10/LQ-Video \
12 --model_path pretrained_models/DOVE \
13 --output_path results/DOVE/UDM10 \
14 --is_vae_st \
15
16# 🔹 Evaluate quantitative metrics
17python eval_metrics.py \
18 --gt datasets/test/UDM10/GT \
19 --pred results/DOVE/UDM10 \
20 --metrics psnr,ssim,lpips,dists,clipiqa💡 If you encounter out-of-memory (OOM) issues, you can enable chunk-based testing by setting the following parameters: tile_size_hw, overlap_hw, chunk_len, and overlap_t.💡 Default save format isyuv444p. If playback fails, trysave_format=yuv420p(may slightly affect metrics).TODO: Add metric computation scripts for FasterVQA, DOVER, and $E^*_{warp}$.











@inproceedings{chen2025dove,
title={DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution},
author={Chen, Zheng and Zou, Zichen and Zhang, Kewei and Su, Xiongfei and Yuan, Xin and Guo, Yong and Zhang, Yulun},
booktitle={NeurIPS},
year={2025}
}