
Vidu S2 treats generated video less like a finished clip and more like a running session: Shengshu Technology's model can alter a live stream while it is playing. ifanr reports that S2-Avatar accepts reference images during livestreaming, allowing instruction-driven changes such as a new prop or background. Its paired S2-Editing model can also modify a playing video stream, including a character subject or visual style.
That changes the job from producing footage to operating an active system. woshipm says S2-Avatar can take fresh instructions and visual references during an interaction, changing what a digital character does next. The useful question therefore includes how convincing the output looks. It also includes how quickly and predictably a live model responds when conditions change.
Video becomes a running session
Vidu S2 frames video generation as an ongoing interaction rather than a clip produced and handed over. Shengshu Technology's system has two parts: S2-Avatar for real-time interaction and S2-Editing for real-time alteration. The distinction matters because the video can keep moving while the operator supplies another instruction or a visual reference.
- 2025Google DeepMind releases Genie 3
- January 2026Synthesia finalizes its Series E financing round
- February 2026Runway completes its financing
With S2-Avatar, a reference image can enter during livestreaming and affect what the digital character does next. ifanr describes uses including prop interaction. It also cites a clothing change or a different background. woshipm further says new instructions and visual references can alter a character's actions during an interaction. They can also change its clothing or the items it uses. The generated character is therefore not treated as fixed when the stream begins.
S2-Editing applies the same idea to an incoming video stream. According to ifanr, one reference image can drive a virtual try-on while video is playing. It can also replace the background, change the style, or replace a character. woshipm says the editing model can work through an open camera or an uploaded existing video.
That changes the unit of work. The relevant output is no longer only a completed video; it is a running session in which later input can revise the next visible moment.
Latency is part of the product
Latency in a live system is the time from a changing condition to a useful correction, not merely the speed of a model's first output. Vidu S2 divides that work between a Backbone network for low-latency streaming of skeletal movement. It also handles physical rules and camera language, while a Refiner network asynchronously rebuilds 720P detail and repairs textures or lighting within milliseconds. That division creates a benchmark question: does the correction arrive before an error becomes visible?
Long-running video adds drift to the test. ifanr reports that Shengshu Technology developed Self-Replay Forcing to address breakdown and drift in extended streaming generation.
Practitioners should change inputs while measuring the whole path, from recognition of the change through instruction handling and rendered response. They should also measure recovery when the result is wrong. Vidu S2's lightweight VLM Agent assesses character status and action progress, then executes a new instruction after the preceding action finishes, according to woshipm. That sequencing can prevent conflicting actions, but it also means an instruction's practical delay depends on the current action. Visual checks must therefore include motion continuity and material texture. They must also determine whether the system returns cleanly after a failed or altered instruction.
Physical guidance raises the cost of a late response. geekpark describes the Strutt ev1c wheelchair as using two lidar sensors and 10 depth sensors. It also uses six ultrasonic sensors plus two cameras for real-time 3D mapping; its Co-Pilot mode avoids pedestrians and pets. It also avoids stair edges. Its Auto-Pilot mode can plan and drive a route after a screen selection or voice command.
For moving equipment, response claims need conditions attached. HeyGo's snowboard sensor samples nine-axis motion signals 100 times per second, while RocXZoom specifies gimbal response as little as 0.01 seconds for targets moving at 45 mph, according to geekpark. Those figures do not establish safe behaviour after ambiguous sensing or a missed target. They also do not establish it after a changed command. Recovery is part of latency.
How live AI systems observe, accept instructions, and change an ongoing task
| Dimension | Vidu S2 | Strutt ev1c | EcoFlow STREAM 5000 |
|---|---|---|---|
| Live task | Real-time video generation and editing | Electric-wheelchair navigation | Home solar power and energy storage |
| What it observes | Character status and action progress are assessed by a lightweight VLM Agent | Real-time 3D mapping from two lidar sensors, 10 depth sensors, six ultrasonic sensors, and two cameras | Dynamic electricity-pricing rules, weather forecasts, and household electricity-use habits |
| How users intervene | New instructions and visual references can change a character's actions, clothing, and items used during an interaction | A user selects a destination on its screen or gives a voice command | Not covered |
| What changes during operation | Style, clothing, character subjects, and backgrounds in a live video stream | It can plan and drive a route after a destination is selected or spoken | Oasis 3.0 can manage electricity while offline in local mode |
| Latency or output detail | 720p output; the Refiner reconstructs 720P details within milliseconds | Not covered | Not covered |
| Failure or safety handling | SRF was developed to address breakdown and drift in long-duration streaming generation | Co-Pilot mode can avoid pedestrians, pets, and stair edges; it also has ABS and an electronic stability system | Not covered |
| Privacy handling | Not covered | Not covered | Not covered |
| Pricing | Not disclosed in sources | Not disclosed in sources | Not disclosed in sources |
A control loop can be an interface
HoverAir Versa shows one path for AI-era hardware: combine familiar devices into a flexible package. According to geekpark, it is presented as both a handheld gimbal camera and a selfie drone. Remove its wings and it works as an Osmo Pocket-type pocket gimbal camera; retain them and the device can transform and deploy within seconds. Its built-in screen, recording button, and joystick make that change legible through ordinary physical controls.
That is integration, rather than a new kind of output.
Generated video points toward a different interface model. Vidu S2 can turn generated or edited footage into synchronized left-eye and right-eye views for compatible headsets, according to ifanr. Existing stereoscopic video can also be edited in real time. The headset is therefore not only a display for a completed clip: it can be part of an editable spatial-video session. woshipm describes Genie 3 as producing explorable worlds at 20 to 24 frames per second, while Runway's Aleph can semantically alter people, objects, backgrounds, camera perspectives, and lighting. Those systems suggest a medium whose visible result remains open to direction during use.
EcoFlow STREAM 5000 sits nearer the hardware pattern, even as its control logic becomes more active. geekpark describes it as a plug-in home solar power and energy-storage system for renters and apartment residents in Europe. Its Oasis 3.0 system can interpret dynamic electricity-pricing rules, view weather forecasts, and learn household electricity-use habits. Here, the product combines energy equipment with adaptive management. With Vidu S2, the adaptation reaches the content itself: the control loop becomes an interface through which a user changes what the audience sees next.
Cheap models do not make live operation cheap
Vidu S2 arrived 69 days after Vidu S1, but the business question is not simply how quickly a new model can improve. S2-Avatar moves real-time output from 540P to 720P, and its two-stage design separates coarse generation from finish work. The Backbone handles low-resolution structure and semantics. It also tracks motion trends, while the Refiner asynchronously adds texture and facial detail.
That split can make a live product feel more capable without making every operating cost disappear. A system serving an active session must support the work happening now, including the capacity demanded when requests arrive together. Higher apparent quality also depends on stages completing in a usable sequence, rather than on a single model benchmark.
Concurrency is a commercial unit, not an abstraction.
According to woshipm, business AI API or Token billing can charge by call volume or concurrency. That matters because a cheaper underlying model does not automatically make a customer-facing live service cheap at busy moments. The bill follows demand for simultaneous work. The operational burden also cannot be reduced to model pricing: people covering exceptions and service failures remain outside a simple per-call calculation.
The woshipm author describes four mainstream AI business models. These include consumer subscriptions and business API or Token billing. They also include project-based or private deployments, along with AI-native tools with value-added services. The same author argues that buyers now fund cost reduction, efficiency improvement, revenue growth, and compliance, rather than AI as an abstract concept. For a live operator, value must support pricing. Repeat purchases must follow as well. Better output matters only if the service can sustain its promised response when customers are actually present.
Choose live AI by the cost of being wrong mid-task
- You need a digital presenter or character to react to new prompts, props, clothing, or backgrounds during a livestream or conversation. Use Vidu S2-Avatar when midstream control matters more than a fixed, pre-rendered clip. It supports reference-image input during livestreaming and can change a character's actions, clothing, and items used; its real-time output is 720p. Test the precise interaction you need, since the sources describe demonstrations but do not disclose reliability, latency, or access conditions.
- You need to alter an incoming camera feed-for example with virtual try-ons, character replacement, style rendering, or background replacement-without interrupting the live stream. Consider Vidu S2-Editing. It supports these changes on a playing video stream and can work through an open camera or uploaded existing video. Use it for reversible visual transformations; do not assume it is appropriate for high-stakes identity or safety decisions, because the source material does not cover error rates, safeguards, or consent controls.
- A physical device must act around people, pets, moving subjects, or hazards. Prioritize sensing, fallback behavior, and manual control over AI autonomy claims. Strutt ev1c has two lidar sensors, 10 depth sensors, six ultrasonic sensors, and two cameras for real-time 3D mapping, while its Co-Pilot and Auto-Pilot capabilities are described as claims. RocXZoom's official specifications state a gimbal response of as little as 0.01 seconds and tracking for targets traveling at 45 mph, but these specifications do not substitute for testing in the actual environment.
- You are selecting an AI product for an enterprise workflow with recurring operational use. Start with a measurable commercial model rather than an abstract AI feature list: consumer subscriptions for feature or computing-power access; API or Token billing for call volume or concurrency; project-based or private deployments for customized local deployment; and AI-native tools with paid advanced features or industry plugins. Use private deployment where local deployment is a requirement; the source material does not say which specific products offer it.
- An AI system must continue operating when connectivity is unavailable or data should remain closer to the user. Favor explicitly local operation where it exists. EcoFlow STREAM 5000 was scheduled to add a local mode in October that can manage electricity while offline. Do not infer equivalent offline capability for Vidu S2, office AI products, or other devices: it is not disclosed in sources.
Put humans around the live system
A bounded live task needs a visible operating contract before it needs a clever model. Show what the system has observed and what it proposes to do next. State which action will occur only after approval. Give the operator an immediate override path. Then retain an audit record of the prompt and sensor inputs. The record should also preserve the proposed action. It should preserve approval and the final result. When an interpretation is wrong, responsibility cannot disappear into the model's output: a named person or team must own the correction and any downstream consequence.
Permission is part of that contract. Vidu S2 lets users converse with a created character after microphone and camera permissions are enabled, according to ifanr. Its VLM Agent also coordinates newly added reference images, ifanr reports. Those inputs should be plainly indicated while active, with a practical way to stop them. A system that can observe a room or hear a conversation needs a privacy check before activation, not a buried setting afterward.
The same test applies beyond video. HeyGo's snowboard-mounted sensor uses nine-axis motion sensing and samples signals 100 times per second, according to geekpark. At that rate, operators need to know which signal triggered an intervention and how to reject it. UTONE, described by geekpark as an always-worn sound-sensing wristband still in development, makes the access question more persistent: when is listening active? Who can review captured data? How is it disabled?
Offline operation also changes the failure plan. EcoFlow STREAM 5000 was scheduled to add a local mode in October for offline electricity management, geekpark reports. Its Oasis 3.0 system can interpret changing electricity-pricing rules. It can view weather forecasts and learn household use habits. Practitioners should define what the system may continue doing when connectivity fails. They should also define what requires confirmation and specify what record remains available for review. Private deployments can pair customized development with local deployment for enterprise and finance customers, as woshipm notes, but local control does not remove the need for accountable human control.
Put a live operator into a bounded task with a clear approval point.
That task may be an energy schedule. It may instead be a visual presentation or coaching feedback. Measure cost reduction or efficiency improvement. Revenue growth and compliance are also outcomes enterprises pay for, according to woshipm.
Then examine the commercial fit. API or Token billing tied to call volume or concurrency can punish a system that stays active; a project-based private deployment may better suit government work. Enterprise and finance work may also fit that model. Consumer subscriptions and free AI-native tools with paid advanced features create different expectations around access and support.
Watch for repeat use, not a compelling demonstration. woshipm argues that value must support pricing. Repeat purchases complete the business foundation; if the operator cannot justify each element inside an existing handoff, it remains an experiment.
For readers outside China
- Availability: Availability outside China is uneven. HoverAir Versa was crowdfunding and scheduled to launch in the United States in autumn 2026. EcoFlow STREAM 5000 is designed for renters and apartment residents in Europe. For Vidu S2, Strutt ev1c, HeyGo, RocXZoom, UTONE, Doubao Work, Qianwen Office, and WorkBuddy, availability by country is not disclosed in sources. HeyGo was recruiting 30 beta testers, while its official release date had not been announced; UTONE had not gone on sale.
- Pricing: The only stated product price is RocXZoom's $449 super early-bird price. Its Kickstarter campaign had a $9,999 target and had raised $247,000. HeyGo's price had not been announced. Pricing for Vidu S2, Strutt ev1c, HoverAir Versa, EcoFlow STREAM 5000, UTONE, and the office AI products is not disclosed in sources. The source material describes consumer AI subscriptions as monthly or annual charges, API or Token billing as charges based on call volume or concurrency, and AI-native tools as free acquisition with charges for advanced features or industry plugins, but gives no product-specific rates.
- Closest Western equivalents: For AI video avatars and generated presenters, HeyGen and Synthesia are the closest named Western comparators.; For semantic video editing of people, objects, backgrounds, camera perspectives, and lighting, Runway's Aleph is the closest named comparator.; For real-time explorable generated worlds, Google DeepMind's Genie 3 is the closest named comparator.
- Data residency: Data-residency policies for Vidu S2, the named Western video products, and the Chinese office AI products are not disclosed in sources. The source material does identify project-based and private deployments as offerings that combine customized development with local deployment for government, enterprise, finance, and similar customers, but does not identify which named product provides that option. EcoFlow STREAM 5000 was scheduled to add an offline local mode in October; this is an operating-mode detail, not a disclosed data-residency policy.
Sources
- ifanr 我用 Vidu S2 找来了「乔布斯」,跟他聊了聊 iPhone Duo https://ifanr.com/1680700
- woshipm ZPedia|视频正在从文件变成界面,Vidu S2 把 AI 视频推向实时交互 https://woshipm.com/ai/6465909.html
- woshipm 拆解AI产品商业模式:别迷信技术,这才是真相! https://woshipm.com/ai/6385587.html
- geekpark 造物 100 #06|自动驾驶上轮椅了,口袋相机学会飞行,AI 教练上了雪场 https://geekpark.net/news/370325
The evidence: 49 facts from 4 Chinese articles
Each line below was extracted from the article it sits under, in Chinese, before any of this was written. The writing is done from these and never from the source prose - that separation is structural, not a promise. How we work.
geekpark造物 100 #06|自动驾驶上轮椅了,口袋相机学会飞行,AI 教练上了雪场
- Apple released its first foldable iPhone Duo on September 10 after years of preparation.
- HoverAir Versa can function as an Osmo Pocket-type pocket gimbal camera after its wings are removed.
- HoverAir Versa has a built-in screen, recording button, and joystick, and can transform and deploy within seconds.
- HoverAir Versa was crowdfunding and was scheduled to launch in the United States in autumn 2026.
- Zero Zero Robotics previously developed the HoverAir X1, an early product in the selfie-drone category.
- Strutt ev1c is an electric wheelchair with two lidar sensors, 10 depth sensors, six ultrasonic sensors, and two cameras for real-time 3D mapping.
- Strutt ev1c has four-wheel independent suspension, ABS, and an electronic stability system, and can be disassembled into five modules for transport in a trunk.
- EcoFlow STREAM 5000 is a plug-in home solar power and energy-storage system designed for renters and apartment residents in Europe.
- EcoFlow STREAM 5000 was scheduled to add a local mode in October that can manage electricity while offline.
- HeyGo's snowboard-mounted motion sensor uses nine-axis motion sensing and samples signals 100 times per second.
- HeyGo was recruiting 30 beta testers, while its official release date and price had not been announced.
- RocXZoom has 20x optical zoom and hybrid zoom equivalent to 2000 mm.
- RocXZoom weighs 400 g, has a 5-inch screen, carries an IP54 dust- and water-resistance rating, and has 4 hours of battery life.
- RocXZoom's Kickstarter campaign had a $9,999 target and had raised $247,000, exceeding its target by 24 times; the super early-bird price was $449.
- UTONE is an always-worn sound-sensing wristband that is still in development and had not gone on sale.
ifanr我用 Vidu S2 找来了「乔布斯」,跟他聊了聊 iPhone Duo
- Shengshu Technology released Vidu S2, a real-time video generation and editing model.
- Vidu S2 consists of S2-Avatar, a real-time interactive model, and S2-Editing, a real-time editing model.
- S2-Avatar supports 720p high-definition output for real-time digital-human interactions.
- S2-Avatar supports reference-image input during livestreaming to enable prop interactions, clothing changes, and background changes based on instructions.
- S2-Editing supports real-time changes to style, clothing, character subjects, and backgrounds in a live video stream.
- Users can upload photos of real people, anime characters, or pets to create a digital character in Vidu S2.
- Users can converse with a created Vidu S2 character after enabling microphone and camera permissions.
- APPSO created a Steve Jobs digital human in Vidu S2 using a full-body reference image and a character description covering personality, speaking style, and attitudes toward products, design, and user experience.
- When asked whether iPhone Duo met his expectations, the Steve Jobs digital human replied, "iPhone Duo is smoother than I expected, but its price makes people hesitate."
- APPSO tested Vidu S2 by directing the Steve Jobs digital human to dance as a test of full-body motion control.
- APPSO tested video calls with a digital character named Himari and a U.S. delivery driver character in his 30s named Tom Miller.
- Vidu S2-Editing can use one reference image to perform real-time virtual try-ons, background replacement, style changes, and character replacement on a playing video stream.
- APPSO obtained real-time try-on results for a blue POLO shirt, a beige coat, a black leather jacket, and a baseball jacket using clothing reference images.
- Vidu S2 uses a two-stage Backbone-Refiner architecture.
woshipmZPedia|视频正在从文件变成界面,Vidu S2 把 AI 视频推向实时交互
- Google DeepMind released Genie 3 in 2025.
- HeyGen disclosed in 2026 that its ARR exceeded $200 million.
- Synthesia's ARR exceeded $100 million in 2025.
- Vidu S2 was released 69 days after Vidu S1.
- Vidu S2 includes S2-Avatar and S2-Editing.
- S2-Avatar increases real-time output resolution from Vidu S1's 540P to 720P.
- Zhang Jintao, a doctoral student of Professor Zhu Jun, led end-to-end research and development for Vidu S2.
- Vidu S2 uses a two-stage architecture with a Backbone network and a Refiner network.
- In January 2026, Synthesia finalized a $200 million Series E financing round at a $4 billion valuation.
- In February 2026, Runway completed $315 million in financing at a post-money valuation of $5.3 billion.
- Tavus completed a $40 million Series B financing round.
woshipm拆解AI产品商业模式:别迷信技术,这才是真相!
- ByteDance fully incorporated the Feishu product team, TRAE, and Coze team into Doubao Work.
- Alibaba incorporated QoderWork, Wukong, and MuleRun into Qianwen Office (千问办公).
- Tencent's WorkBuddy is positioned as an all-scenario AI office workbench.
- WorkBuddy brought QQ Show and QQ Pet-style gameplay into office desktops.
- WorkBuddy launched Peacekeeper Elite co-branded IP skins, a growth plan, blind-box desktop pets called Buddy Cat, and a travel-assignment development mechanism.
- Consumer AI subscriptions charge monthly or annually to unlock features or computing power, with ChatGPT Plus given as an example.
- Business AI API or Token billing charges based on call volume or concurrency.
- Project-based AI offerings and private deployments combine customized development with local deployment for government, enterprise, finance, and similar customers.
- AI-native tools with value-added services acquire users for free and charge for advanced features or industry plugins.