Views
No views yet
Imagine How to Change: Explicit Procedure Modeling for Change Captioning
5h7w.h5 files) generated by VFIformer.CLEVR-dataedit-dataspot-datafiltered-spot-captionspretrained_vqgan – VQGAN models for each datasetstage1_clevr_beststage1_edit_beststage1_spot_bestclevr_bestedit_bestspot_bestNote: Stage 1 checkpoints can be directly reused to initialize Stage 2 training.
densevid_evalfiltered-spot-captions to the original caption directory of the Spot-the-Diff dataset.| Dataset | Folder | Rename To |
|---|---|---|
| CLEVR-Change | CLEVR-data | CLEVR_processed |
| Image-Editing-Request | edit-data | edit_processed |
| Spot-the-Diff | spot-data | spot_processed |
filter_files in the project root directory.pretrained_vqgan in the project root directory.symlink_path in training scripts as:symlink_path="/path/to/stage1/weight/dalle.pt"resume_path in evaluation scripts as:resume_path="/path/to/pretrained/model/model.chkpt"densevid_eval directory in the project root before evaluation.1@inproceedings{
2 sun2026imagine,
3 title={Imagine How To Change: Explicit Procedure Modeling for Change Captioning},
4 author={Sun, Jiayang and Guo, Zixin and Cao, Min and Zhu, Guibo and Laaksonen, Jorma},
5 booktitle={The Fourteenth International Conference on Learning Representations},
6 year={2026},
7}