: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning
FlowAct-R2 introduces a Streaming Multimodal Reference DiT for streaming digital human video generation, built by adapting the pretrained Seedance 2.0 Mini reference-to-video backbone. By extending Diffusion Forcing to streaming multimodal references, the model continuously incorporates rolling action prompts, audio, and dynamically updated image, audio, and video references while preserving bidirectional spatiotemporal modeling. A video-driven RoPE aligns incoming reference chunks with the generation timeline, enabling references to be updated or switched during an ongoing stream. FlowAct-R2 further combines explicitly trained reference-plus-image conditioning with partially noised historical motion frames to preserve identity and appearance while reducing error accumulation.
Complementing the video generator, the Proactive Interaction Agent enables a digital human to proactively sustain a live session while remaining responsive to audience interactions. Offline, the agent derives a persona from multimodal character assets, organizes the session into a long-horizon agenda, and compiles complex behaviors into reusable multimodal skills. Online, it combines the prepared agenda and skills with audience events and summarized interaction history to determine what behavior to perform and how it should be scheduled. By initiating, overlaying, deferring, interrupting, or resuming behaviors according to the current context, the digital human can continue purposeful activities without continuous user prompts while responding naturally when new interactions arise.
FlowAct-R2 simulates lifelike entertainment streamers that interact with live audience messages and invoke skills to perform singing, dancing, and other talent-driven behaviors.
With streaming multi-reference conditioning, FlowAct-R2 enables live shopping avatars to hold products, try on items, and respond to audience feedback in interactive product-selling streams.
FlowAct-R2 enables natural video chatting with responsive listening, expressive turn-taking, and lifelike avatar behaviors for seamless real-time interaction.
FlowAct-R2 supports vlog streams with multi-scene switching and audience-driven branching, allowing live viewer inputs to steer the next scenario and action.
Enabled by the planning capabilities of its proactive agent, FlowAct-R2 delivers clearer videos, more natural and expressive motions, stronger action responsiveness, and more accurate, contextually appropriate dialogue than Vidu-S1, achieving GSB scores of +54.76% for video quality and +40.48% for real-time interaction in human evaluations, whereas Vidu-S1 is comparatively static and may introduce unrelated conversational memory.
Compared with Vidu-S2, FlowAct-R2 better preserves the appearance of referenced objects and switches more naturally between product references, demonstrating stronger object consistency and reference switching in these live shopping examples. Both Vidu-S2 and FlowAct-R2 can trigger dance from text; FlowAct-R2 further supports reference-video conditioning, enabling more expressive dance generation.
@online{2026flowact-R2,
title={FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning},
author={Ziyao Huang, Zhengkun Rong, Shiyang Qin, Shuang Liang, Wentao Hu, Yuxuan Luo, Yuan Zhang, Mingyuan Gao},
year={2026},
url={https://bone-11.github.io/Flowact-R2/},
}
@article{wang2026flowact,
title={FlowAct-R1: Towards Interactive Humanoid Video Generation},
author={Wang, Lizhen and Zhu, Yongming and Ge, Zhipeng and Zheng, Youwei and Zhang, Longhao and Hu, Tianshu and Qin, Shiyang and Luo, Mingshuang and Zhang, Jiaxu and Chen, Xin and others},
journal={arXiv preprint arXiv:2601.10103},
year={2026}
}
@article{zhu2024infp,
title={INFP: Audio-driven interactive head generation in dyadic conversations},
author={Zhu, Yongming and Zhang, Longhao and Rong, Zhengkun and Hu, Tianshu and Liang, Shuang and Ge, Zhipeng},
journal={arXiv preprint arXiv:2412.04037},
year={2024}
}