FlowAct-R2 Logo : Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning

Bytedance Intelligent Creation
*Core Contributors Corresponding Author
FlowAct-R2 Overview

We present FlowAct-R2, a next-generation framework for proactive, multimodal, and highly interactive humanoid video generation, enabling intelligent real-time streaming with rich behaviors across diverse live scenarios.

  • Streaming Multimodal Reference Generation: By extending diffusion forcing from chunkwise generation to streaming multimodal references, FlowAct-R2 enables real-time video generation with rolling prompts and multimodal references, including continuously updated images, audio, and video, and supports hour‑level ultra‑long video generation with 720p.
  • Proactive Interaction Agent: To enable digital humans to autonomously drive their behavior when no external input is present, while seamlessly responding when interactions arise, FlowAct-R2 introduce an agent workflow with pre-online planning and online scheduling and response, which supports persona-driven behaviors, interruption handling, and skill execution for autonomous live interactions.
  • Rich Behaviors Across Diverse Live Scenarios: FlowAct-R2 supports real-time text and audio interactions across four representative scenarios—video chatting, live shopping, entertainment streaming, and interactive gaming—generating not only conversational responses but also expressive behaviors and task-oriented actions beyond conventional question answering.

Method Overview

Interpolate start reference image.

FlowAct-R2 introduces a Streaming Multimodal Reference DiT for streaming digital human video generation, built by adapting the pretrained Seedance 2.0 Mini reference-to-video backbone. By extending Diffusion Forcing to streaming multimodal references, the model continuously incorporates rolling action prompts, audio, and dynamically updated image, audio, and video references while preserving bidirectional spatiotemporal modeling. A video-driven RoPE aligns incoming reference chunks with the generation timeline, enabling references to be updated or switched during an ongoing stream. FlowAct-R2 further combines explicitly trained reference-plus-image conditioning with partially noised historical motion frames to preserve identity and appearance while reducing error accumulation.

Complementing the video generator, the Proactive Interaction Agent enables a digital human to proactively sustain a live session while remaining responsive to audience interactions. Offline, the agent derives a persona from multimodal character assets, organizes the session into a long-horizon agenda, and compiles complex behaviors into reusable multimodal skills. Online, it combines the prepared agenda and skills with audience events and summarized interaction history to determine what behavior to perform and how it should be scheduled. By initiating, overlaying, deferring, interrupting, or resuming behaviors according to the current context, the digital human can continue purposeful activities without continuous user prompts while responding naturally when new interactions arise.

Entertainment Streaming

FlowAct-R2 simulates lifelike entertainment streamers that interact with live audience messages and invoke skills to perform singing, dancing, and other talent-driven behaviors.

Live Shopping

With streaming multi-reference conditioning, FlowAct-R2 enables live shopping avatars to hold products, try on items, and respond to audience feedback in interactive product-selling streams.

Video Chatting

FlowAct-R2 enables natural video chatting with responsive listening, expressive turn-taking, and lifelike avatar behaviors for seamless real-time interaction.

Live Vlogging

FlowAct-R2 supports vlog streams with multi-scene switching and audience-driven branching, allowing live viewer inputs to steer the next scenario and action.

Comparing to Vidu-S1

Enabled by the planning capabilities of its proactive agent, FlowAct-R2 delivers clearer videos, more natural and expressive motions, stronger action responsiveness, and more accurate, contextually appropriate dialogue than Vidu-S1, achieving GSB scores of +54.76% for video quality and +40.48% for real-time interaction in human evaluations, whereas Vidu-S1 is comparatively static and may introduce unrelated conversational memory.

Comparing to Vidu-S2

Compared with Vidu-S2, FlowAct-R2 better preserves the appearance of referenced objects and switches more naturally between product references, demonstrating stronger object consistency and reference switching in these live shopping examples. Both Vidu-S2 and FlowAct-R2 can trigger dance from text; FlowAct-R2 further supports reference-video conditioning, enabling more expressive dance generation.

BibTeX


@online{2026flowact-R2,
  title={FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning},
  author={Ziyao Huang, Zhengkun Rong, Shiyang Qin, Shuang Liang, Wentao Hu, Yuxuan Luo, Yuan Zhang, Mingyuan Gao},
  year={2026},
  url={https://bone-11.github.io/Flowact-R2/},
}      
      
@article{wang2026flowact,
  title={FlowAct-R1: Towards Interactive Humanoid Video Generation},
  author={Wang, Lizhen and Zhu, Yongming and Ge, Zhipeng and Zheng, Youwei and Zhang, Longhao and Hu, Tianshu and Qin, Shiyang and Luo, Mingshuang and Zhang, Jiaxu and Chen, Xin and others},
  journal={arXiv preprint arXiv:2601.10103},
  year={2026}
}

@article{zhu2024infp,
  title={INFP: Audio-driven interactive head generation in dyadic conversations},
  author={Zhu, Yongming and Zhang, Longhao and Rong, Zhengkun and Hu, Tianshu and Liang, Shuang and Ge, Zhipeng},
  journal={arXiv preprint arXiv:2412.04037},
  year={2024}
}