Shengshu Tech launches the Vidu S1 real-time interaction model at the 2026 Global Digital Economy Conference.

iconKuCoinFlash
Share
AI summary iconSummary
Shengshu Tech unveiled Vidu S1, a real-time interaction model, at the 2026 Global Digital Economy Conference. The model supports voice control, 540p resolution, and 42 FPS, enabling users to generate interactive digital characters from a single image. Zhu Jun, founder of Shengshu Tech, presented the model at the AI Fusion Application Development Forum. The company was also named a "New Model and Application Benchmark Enterprise" in the 2025 Beijing Digital Economy Benchmark Enterprise Evaluation Report. Digital asset news highlights the growing role of AI in the digital collectibles sector.
ME AI News, on July 3, at the Forum on the Development of AI Integration Applications during the 2026 Global Digital Economy Conference, Zhu Jun, founder of Shengshu Technology, delivered a keynote speech titled “General World Models: A New Paradigm for Unifying the Digital and Physical Worlds,” and officially launched Vidu S1, the next-generation model designed for real-time interactive scenarios. During the conference, the Beijing Association of Software and Information Services (BSIA) officially released the “2025 Beijing Digital Economy Benchmark Enterprise Evaluation Report,” and Shengshu Technology was successfully selected as a “Benchmark Enterprise for New Models and New Applications” for its outstanding performance in technological innovation and industrial application. The Vidu S1 real-time interactive model delivers a new generation of video generation capabilities that are real-time and interactive, advancing AI video from “generating a single clip” to “sustained interaction.” The model supports real-time video calls and voice-controlled video direction, enabling users not only to control digital characters through voice commands but also to achieve unlimited-duration continuous interaction. Meanwhile, Vidu S1 supports 540P (960×540) high-definition resolution and 25 FPS frame rate (up to 42 FPS), allowing rapid creation of personalized interactive characters based on any initial avatar—real human, anime, or pet—and customized voice tones, delivering users a more natural, fluid, and immersive real-time interactive experience. Voice commands are followed in real time—transitioning from offline generation to real-time response, enabling digital characters to truly “understand” users. Traditional large video models typically operate in an offline mode: “input prompt → wait for generation → play result.” Once generated, the content and trajectory are largely fixed; any adjustment to actions or plot requires re-inputting prompts and regenerating—maintaining an offline “generate-and-view” relationship between user and video. Vidu S1 breaks through this boundary. Users can continuously input voice commands during a video call, and the model combines the voice content, conversational context, and current visual state to generate subsequent character actions and content in real time. Moreover, Vidu S1 elevates digital characters from “voice-driven lip-sync” to “voice-controlled behavior.” Unlike traditional digital characters that rely on “audio-driven lip-sync + pre-set action libraries,” Vidu S1 employs real-time video generation technology to transform voice from an audio signal driving lip movement into real-time instructions controlling visual behavior. The model not only generates synchronized lip movements but also understands semantics, intent, and emotion to produce matching facial expressions, eye gaze, gestures, body posture, and full-body motions in real time—transforming digital characters from “talking avatars” into generative agents capable of understanding users, responding instantly, and sustaining continuous interaction. Unlimited-duration real-time generation enables videos to evolve continuously through interaction. Traditional video generation models typically produce fixed-length clips of 3–30 seconds at once. During generation, users cannot insert new commands or alter subsequent visuals in real time. Vidu S1 adopts an autoregressive diffusion (AR + Diffusion) approach—not generating the entire video at once, but continuously predicting and generating subsequent content based on previously generated frames, current voice commands, and conversational context. When users issue new voice instructions, the model instantly interprets them and adjusts the character’s expressions, actions, and future video trajectory—transforming pre-determined static content into a continuously generated, real-time responsive, dynamically evolving interactive process. Beyond interactive real-time generation, Vidu S1 also achieves unlimited-duration real-time video generation for the first time. Even after continuous generation for hours, visual quality remains stable without drift or collapse. Achieving long-term continuous interaction requires more than just “continuous generation.” The model must simultaneously maintain character identity stability, natural and coherent motion, continuous reception of user commands, and real-time responsiveness throughout extended operation. Vidu S1 sustains stable character appearance and fluid motion over long durations while continuously accepting voice commands and responding in real time—pioneering unlimited-duration generative video interaction. Customizable characters—no modeling or training required; one image creates a real-time interactive character. Traditional digital character creation typically requires uploading multiple images or video clips followed by modeling, rigging, lip-sync adaptation, and separate training—a lengthy process. Vidu S1 employs a purely generative approach—no offline modeling or training is needed for each character. Users simply upload one initial image; the model instantly understands the character’s identity, appearance, and visual style, then generates synchronized lip movements, expressions, actions, and body posture in real time. Whether real human, anime character, or pet image, any subject can be rapidly transformed into a real-time interactive generative character. Vidu S1 also supports customizable voice tones to unify visual identity with auditory identity. The process of creating digital characters shifts from “upload assets → wait for training” to “upload image → interact immediately,” dramatically lowering the barrier to creating personalized real-time characters. 540P · 25FPS Real-Time Interaction: Delivering Video Call-Quality Experience Real-time interaction demands not only streaming generation but also high resolution and frame rate during real-time output. Vidu S1 is optimized for real-time interactive scenarios through coordinated improvements in model acceleration, inference engine design, and cluster deployment strategies—achieving 540P (960×540) high-definition resolution and 25 FPS (up to 42 FPS) smooth frame rates for real-time video generation. On the model side, Vidu S1 leverages Shengshu Technology’s TurboDiffusion [1] inference acceleration framework using techniques such as few-step generation, low-bit attention SageAttention [2], sparse attention SLA [3], and SpargeAttention [4] to significantly reduce computational cost per frame—enabling 540P resolution at 25 FPS (up to 42 FPS) real-time generation even on consumer-grade GPUs. On the system side, Vidu S1 utilizes Shengshu Technology’s TurboServe [5] inference deployment engine for efficient request scheduling. The system continuously records user inputs, character states, and historical frames, dynamically allocating computational resources based on interaction status. Through coordinated optimization of model inference and streaming services, Vidu S1 achieves a critical leap—from “generating videos faster” to “keeping videos continuously online, stably outputting, and responding in real time.” 540P resolution and 25 FPS (up to 42 FPS) are not merely metrics—they signify that real-time video generation now possesses the foundational technical capability to enter applications such as video calling, interactive live streaming, real-time companionship, interactive gaming, and XR. As large video models continue evolving, industry competition is shifting from isolated capabilities like resolution, duration, and speed toward a holistic competition centered on real-time performance, controllability, and interactivity. The launch of Vidu S1 transforms video from static, pre-generated content meant for offline viewing into an interactive medium capable of understanding commands, responding in real time, and evolving continuously.In the future, Vidu S1 can be widely applied in scenarios such as AI emotional companionship, AI virtual idols, interactive live streaming, game NPCs, brand digital humans, intelligent customer service, online education, and XR, transforming digital characters from one-time content assets into persistent, continuously interactive intelligent gateways. From generating a single video to creating a character capable of ongoing interaction—from offline content output to real-time two-way communication—Vidu S1 further expands the boundaries of video large models, ushering AI video generation into a new era of real-time interactivity. Vidu S1 is now open for beta testing; users can customize initial images and experience real-time interaction: Online experience: https://www.vidu.cn/vidu-stream API experience: https://platform.vidu.cn/live/landing App experience: Search for "Vidu AI Pro" in your mobile app store to download the latest version, then tap "Vidu S1" within the app to get started. [1] TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times. [2] Sageattention: Accurate 8-bit attention for plug-and-play inference acceleration. [3] Sla: Beyond sparsity in diffusion transformers via fine-tunable sparse-linear attention. [4] Spargeattention: Accurate and training-free sparse attention accelerating any model inference. [5] TurboServe: Serving Streaming Video Generation Efficiently and Economically. (Source: Ifnar)
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.