As AI applications continue to advance, digital humans are gradually evolving from showcase-style content toward real-time interactive and always-online application formats. In line with this trend, the Soul Zhang Lu team has been steadily driving the production-grade deployment of multimodal generation technology. Recently, its AI team (Soul AI Lab) released the open-source model SoulX-LiveAct. Targeting the “long-duration instability” problem in real-time digital human generation, the model proposes a new solution approach, achieving a fresh balance between stability and efficiency in real-time digital human video generation.
Currently, digital human live streaming, video podcasts, and real-time interactive scenarios impose increasingly demanding requirements on generation models. Traditional methods tend to exhibit identity drift, detail loss, and frame flickering as generation duration extends, while inference costs grow over time. SoulX-LiveAct is engineered around these issues, introducing two core mechanisms—Neighbor Forcing and ConvKV Memory—that enable autoregressive diffusion (AR Diffusion) to maintain consistency and controllability throughout extended generation.

Architecturally, SoulX-LiveAct employs a chunk-based autoregressive generation approach, splitting video into consecutive frame blocks for processing. Within each chunk, the model captures detail through the diffusion process, while contextual information is passed between chunks to ensure seamless continuity of identity and motion. This mechanism provides the structural foundation for streaming generation while supporting long-duration consistency.
The Neighbor Forcing mechanism further optimizes temporal sequence modeling on top of this foundation. Unlike conventional autoregressive methods that pass states across different diffusion steps, this mechanism aligns adjacent-frame latents within the same diffusion step, enabling the model to propagate conditions in a unified noise semantic space. This reduces distributional discrepancies between training and inference, helping the model maintain stable temporal relationships during long-sequence generation and mitigating the impact of error accumulation.
For memory management, SoulX-LiveAct introduces the ConvKV Memory mechanism, transforming the traditionally linearly growing attention cache into a “short-term precise + long-term compressed” structure. Recent information is retained through a high-precision window to ensure continuity of local details, while older information is compressed via lightweight convolution into fixed-length historical context representations. This design enables constant GPU memory usage during extended generation, preventing resource consumption from climbing endlessly as historical information accumulates. Additionally, a RoPE Reset mechanism aligns position encodings, effectively mitigating potential positional drift in long sequences.

In terms of inference efficiency, SoulX-LiveAct prioritizes stable latency and controllable resource usage. By transforming historical context into a fixed-budget memory structure, the model achieves constant GPU memory consumption during inference, avoiding performance degradation as generation duration increases. At 512×512 resolution, the model achieves 20 FPS streaming inference on 2 H100/H200 GPUs, with end-to-end latency of approximately 0.94 seconds and a computational cost of roughly 27.2 TFLOPs/frame. This performance provides a viable engineering foundation for sustained online applications.
In terms of evaluation results, SoulX-LiveAct demonstrates stable performance across multiple benchmarks. On the HDTF dataset, it achieves a Sync-C of 9.40 and Sync-D of 6.76, with distribution similarity scores of 10.05 FID and 69.43 FVD. In VBench, Temporal Quality reaches 97.6 and Image Quality 63.0, with VBench-2.0 Human Fidelity at 99.9. On the EMTD dataset, the model maintains strong synchronization and quality, with Sync-C of 8.61 and Sync-D of 7.29. VBench Temporal Quality is 97.3, Image Quality 65.7, and Human Fidelity 98.9—reflecting the model’s stable capabilities in complex motion and full-body expression scenarios.

Building on these capabilities, SoulX-LiveAct is applicable to a wide variety of scenarios requiring sustained online operation, including digital human live streaming, AI-powered education, intelligent customer service, premium content delivery, and podcast production. In open-world interactive environments, where digital characters must continuously maintain consistency in speech, motion, and identity, the model’s performance in full-body motion and facial expression provides the technical underpinning for such applications.
In recent years, the Soul AI team has consistently pursued an open-source strategy, successively releasing models such as SoulX-FlashTalk and SoulX-FlashHead. Beyond real-time digital human generation, the team has also expanded into multimodal directions, launching the podcast speech synthesis model SoulX-Podcast, the singing voice synthesis model SoulX-Singer, and the full-duplex voice interaction module SoulX-Duplug—progressively building a technology ecosystem centered on real-time interaction.
By open-sourcing SoulX-LiveAct, the Soul Zhang Lu team further refines the technical path between long-duration stability and real-time inference in the real-time digital human generation space. Through its ongoing commitment to open source, Soul’s technologies serve not only its own product ecosystem but also provide the industry with referenceable implementation approaches, driving digital human applications toward greater stability and practical viability.

