CoInteract is the first end-to-end framework that generates physically-consistent human-object interaction (HOI) videos with zero additional inference cost. Given a person image, a product image, text prompts, and optional speech audio, CoInteract produces realistic videos where humans naturally grasp, wear, present, and manipulate objects — with no hand-object interpenetration or geometric misalignment.

| Category | Examples |
|---|---|
| 🤲 Grasping | Macaron box · Teapot · Skincare serum · Coffee mug |
| 👜 Presenting | Leather handbag · Eyeshadow palette · Decorative plate |
| 👗 Wearing | Emerald necklace · Sports jacket |
| 🌵 Holding | Cactus pot · Various daily objects |
| Property | Value |
|---|---|
| Backbone | Diffusion Transformer (DiT) |
| Resolution | 720p 480p |
| Frame Count | Up to 81 frames per chunk |
| Multi-Chunk | Supported for long-form video |
| Inputs | Person image + Object image + Text prompt + Audio + (Optional) Pose |
| Training Data | Large-scale HOI video dataset with structure annotations |
| Precision | FP16 / BF16 |
| License | Apache 2.0 |

Full quantitative results will be released upon paper acceptance.
1@article{luo2025cointeract,
2 title={CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation},
3 author={Luo, Xiangyang and Xin, Xiaozhe and Feng, Tao and Guo, Xu and Jin, Meiguang and Ma, Junfeng},
4 journal={arXiv preprint arXiv:2604.19636},
5 year={2026}
6}