94006622e68aead770c906de930643ed60c0b293
Making Avatars Interact
Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
InteractAvatar is a novel dual-stream DiT framework that enables talking avatars to perform Grounded Human-Object Interaction (GHOI). Unlike previous methods restricted to simple gestures, our model can perceive the environment from a static reference image and generate complex, text-guided interactions with objects while maintaining high-fidelity lip synchronization.
ComfyUI_InteractAvatar
InteractAvatar is a novel dual-stream DiT framework that enables talking avatars to perform Grounded Human-Object Interaction (GHOI)
Coming soon
- Need Ram>32G , Vram> 8G (offload mode)
Example
Citation
@article{zhang2026making,
title={Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars},
author={Zhang, Youliang and Zhou, Zhengguang and Yu, Zhentao and Huang, Ziyao and Hu, Teng and Liang, Sen and Zhang, Guozhen and Peng, Ziqiao and Li, Shunkai and Chen, Yi and Zhou, Zixiang and Zhou, Yuan and Lu, Qinglin and Li, Xiu},
journal={arXiv preprint arXiv:2602.01538},
year={2026}
}
🙏 Acknowledgements
We sincerely thank the contributors to the following projects:
Languages
Python
99.6%
Shell
0.4%