
WALL-SS simulates 60-second tasks such as pouring water and organizing objects, then compares the outcomes of different actions before a robot acts. ifanr reports that its team also retained data from missed grasps, slips, collisions, re-grasps, manual takeovers, and recovery attempts.
That pairing defines the emerging test: a virtual world can screen behavior, but field conditions decide whether it is useful. geekpark describes Galbot transferring tennis skills from virtual training to a court shaped by uneven friction, shifting light, wind interference, and equipment noise. Its virtual environment used Yinhe Xingfang to turn imperfect human data into training material and run multi-agent adversarial training. The practical question is how simulation, failure records, and deployment measures can become one validation loop.
Simulation needs a real-world verdict
WALL-SS treats simulation less as a stage for a polished demonstration than as a controlled way to ask what a policy will do next. According to ifanr, it can simulate 60-second tasks such as pouring water and organizing objects, then compare outcomes from different actions. The team retained training records from missed grasps and collisions, along with re-grasps, manual takeovers, and failure recovery. The point is evidence, not spectacle.
- AprilZhiyuan Robot presents seven deployment-oriented productivity solutions
- JuneGenie G2 robots enter Longqi Technology's mass-production factory
- August 22Galbot humanoid plays a tennis exhibition match at the robot games opening
Virtual environments can also create repetitions that hardware cannot easily supply. geekpark reports that Yinhe Xingfang converts imperfect human data into training data and supports multi-agent adversarial training in virtual tennis.
Galbot says its robot learned explosive running and continuous shot control through tens of millions of trial-and-error shots in a virtual environment. But a simulation score only matters when its ordering survives transfer. Galbot states that its skills reached a real tennis court despite uneven ground friction, changing light, wind interference, and equipment noise. WALL-SS supplies a stricter gate: five WALL-WM policy versions were tested across six tasks and 20 matched initial-state groups, in simulation and on real robots. That produced 600 closed-loop result pairs. In 527, final success or failure matched; across 30 task-and-policy-version combinations, the success-rate correlation was 0.926 and mean absolute error was 0.062.
Failure data makes the useful test set
WALL-SS treats error records as training material rather than footage to edit away. According to ifanr, its retained data includes missed grasps, slips, collisions, re-grasps, manual takeovers, and failure recovery. During training, the team perturbs historical information and asks the model to infer the correct future from inputs that already contain errors, using scale-by-scale dream forcing. Recovery becomes part of validation.
In a 60-second water-pouring simulation, frame-by-frame robotic-arm trajectory error stayed below 0.5% of the image diagonal. Yet ifanr says a version retaining only recent segments deteriorated noticeably after 20 to 30 seconds. Online visual-dynamics policy alignment raised action following from 0.264 to 0.290 and trajectory accuracy from 0.512 to 0.539. It also reduced cross-segment boundary error from 0.118 to 0.104.
Those results describe a virtual test, not evidence that every physical edge case has been covered.
The real-world record has a different job. Tang Mu argues in woshipm that commercial coffee robots must handle real orders, customers, payment pressure, equipment failures, replenishment needs, output stability, and single-location ROI. He distinguishes operating data from training data: it first informs product definition, operational improvement, and business decisions. Guo Renjie likewise said his team tested hair, milk, water, cat food, dog food, cat litter, and instant noodles before a Bilibili video of vacuuming instant noodles spread widely. A polished clip can show completion; retained failures show what must be designed, measured, and recovered from.
A tennis court is not a product brief
Galbot's tennis exhibition at the second World Humanoid Robot Games showed a broad physical repertoire, according to geekpark. On August 22, its humanoid served and returned balls, played from the baseline, volleyed, and used forehands and backhands. In doubles, it worked with a human teammate in real time, changing tactics as court conditions changed. geekpark also reports a difficult save after it crossed more than half the court, followed by autonomous recovery from a fall.
That is meaningful evidence of mobility and perception under public performance pressure. It is not, by itself, a product brief.
Zhiyuan Robot's results make a related point. ifanr reports that it led the medal table with 18 gold medals, 16 silver medals, and 12 bronze medals. The company competed at the World Humanoid Robot Games for the first time using mass-produced models rather than competition-customized machines. That makes the result more relevant to repeatability, but medals still describe competitive performance rather than a buyer's operating need.
As woshipm's Tang Mu argues, "humanoid robot" identifies an appearance, not the task a customer needs done. A commercial robot needs a defined use case and market. Coffee robots illustrate the harder test: real orders, customers, payment pressure, equipment failures, replenishment, and stable output. Tang Mu says product managers should measure utilization, MTBF, scene-adaptation speed, consistency, 7x24 maintenance capability, and single-location ROI. Those measures turn an impressive court appearance into a question of delivery: who pays, what the machine must reliably do, and whether the operation can sustain itself.
Embodied-AI validation stack: simulation, field operations, and interaction design

| Virtual training and world models | Commercial deployment and operations | Human-centered interaction and safety | |
|---|---|---|---|
| Primary role | Pre-train skills, simulate outcomes, and screen policy behavior before real-world execution. | Define products, improve operations, and support business decisions through operating data. | Enable people to cooperate with robots intuitively without accommodating or guarding against them. |
| Key technical approach | Galbot uses virtual tennis training and multi-agent adversarial training; WALL-SS puts actions and visual frames on the same timeline. | XBOT tracks operating data, location data, order heat maps, and operating status; Zhiyuan Robot deploys robots in factories. | Use layered processing: language for task-level intent, vision for coordination and prediction, and force sensing for contact coordination. |
| Treatment of failures | WALL-SS training data retains missed grasps, slips, collisions, re-grasps, manual takeovers, and failure recovery; dream forcing predicts futures from error-containing inputs. | Commercial robots face equipment failures, replenishment needs, output-stability requirements, and single-location ROI requirements. | Safety-related tasks require deterministic design and protection against unexpected situations; contact safety cannot rely only on vision, software judgment, and emergency stopping. |
| How behavior is checked | WALL-SS was tested against real robots with closed-loop virtual and real-robot result pairs; Galbot transferred virtual skills to a real court with changing conditions. | Operators assess utilization, MTBF, scene-adaptation speed, output stability, consistency, 7x24 maintenance capability, and ROI. | People learn a robot's behavior through consistent movement responses, preparatory motions, yielding signals, and sensing before releasing a handed-over object. |
| What simulation or data cannot replace | Virtual results require real-robot comparison and transfer across uneven friction, changing light, wind interference, and equipment noise. | Operational data is not equivalent to training data. | A single perception-thinking-action pipeline is insufficient for every task; cameras alone are insufficient for intent recognition. |
| Evidence cited in sources | WALL-SS reported matching final success or failure outcomes for 527 of 600 virtual and real-robot result pairs; Galbot reported tens of millions of virtual trial-and-error shots. | XBOT says its coffee robots have been deployed globally and that commercial operation exposes robots to real orders, customers, payments, failures, and replenishment. | The interaction framework distinguishes extremely fast unconscious channels, medium-speed semi-automatic channels, and slow conscious language. |
Human interaction sets the safety boundary
Guo Renjie treats safety and full automation as prerequisites for consumer robots entering homes. That standard frames the household robot as a system that must be trusted before it is given ordinary domestic responsibility. The interaction-design argument from woshipm sets a more operational boundary: autonomy should vary with the consequence of an action, rather than applying one perception-thinking-action pipeline to every task.
Human interaction is not a language prompt followed by a robot response.
According to woshipm, people coordinate through nonverbal, instantaneous, bidirectional signals as well. A robot therefore needs to make its state legible in motion: slowing down can show that it is yielding, constant movement can indicate normal operation, and stopping can communicate uncertainty and a request to wait. That makes an explicit handoff part of safety. A robot handing over an object should not release it until it senses that the recipient has securely taken it.
The same logic limits where autonomy belongs. woshipm argues that low-cost, reversible actions can proceed autonomously, while costly or irreversible actions need greater confidence or human confirmation. Safety-critical tasks require deterministic design and protection for unexpected situations. Contact safety also cannot depend solely on vision, danger assessment, and an emergency stop, since the response may come after contact. Tactile sensing at the end effector and fast local control can detect a slipping grasp and increase force quickly. The practical goal is selective autonomy with visible safeguards, not an all-or-nothing claim of independence.
Where simulation fits-and where field validation must take over
- You need to screen long-horizon tabletop policies before using scarce real-robot time. Use a world model such as WALL-SS as a pre-screening layer, particularly when action consequences need to remain coherent across a 60-second task. In tests across six tasks and 20 matched initial-state groups, 527 of 600 virtual and real-robot result pairs matched on final success or failure; virtual and real success rates across 30 task-and-policy-version combinations had a correlation coefficient of 0.926. Treat that as evidence for prioritization, not a substitute for physical testing.
- Your robot must cope with missed grasps, slips, collisions, or operator intervention. Build failure and recovery traces into training and validation rather than training only on clean demonstrations. The WALL-SS team retained data covering missed grasps, slips, collisions, re-grasps, manual takeovers, and failure recovery, and trained with error-containing histories. Simulation can expose a policy to these cases repeatedly, but real-world recovery behavior still needs validation on hardware.
- You are building a robot that works close to people, especially in a consumer setting. Do not use a single slow perception-reasoning-action loop for all interaction. Use fast monitoring for movement, approaching stimuli, and contact; use vision for coordination and prediction; and use force sensing for real-time contact coordination. Safety-related functions should remain deterministic and retain protection for unexpected situations. For irreversible or high-cost actions, require higher confidence or human confirmation.
- A demo, competition result, or simulation benchmark looks impressive, and you need to decide whether it is a product. Move the evaluation to operating metrics: utilization, MTBF, scene-adaptation speed, output stability, consistency, 7x24 maintenance capability, and single-location ROI. XBOT describes commercial coffee-robot deployment as exposure to orders, customers, payment pressure, equipment failures, replenishment, and output-stability requirements; it also argues that operating data first serves product definition, operational improvement, and business decisions rather than automatically becoming training data.
- You need to decide whether a behavior should run autonomously in public or commercial operations. Autonomous execution is more appropriate for low-cost, reversible actions. Raise the confidence threshold-or ask for human confirmation-for irreversible, high-cost actions. Make state legible through motion: recognizable preparatory movements, slowing to yield, steady speed for normal operation, and stopping to signal uncertainty can help people form reliable expectations.
Deployment metrics decide the robot form
A robot's form is validated by the job it can sustain, not by the label attached to its body. According to woshipm, Tang Mu argues that "humanoid robot" describes appearance rather than the user need or task being solved. That distinction changes deployment review from a demo judgement into an operating test.
Zhiyuan Robot's Genie G2 offers one factory-side reference point. ifanr reports that the model had been deployed at scale in Longqi and SAIC factories before the games. In June, multiple Genie G2 robots entered Longqi Technology's mass-production factory in Nanchang, Jiangxi, where they completed a cumulative 17,625 operations at a 99.99% task success rate. Those figures matter because they connect a robot to repeated work under production conditions.
A high task-success figure is a starting point, not the whole verdict.
ifanr also describes seven deployment-oriented productivity solutions presented in April, spanning production-line loading and unloading, industrial handling, logistics sorting, guided tours and shopping assistance, service retail stations, security patrols, and industrial and commercial cleaning. Each setting creates a different operating burden. Practitioners need to track utilization, MTBF, scene-adaptation speed, output stability, consistency, 7x24 maintenance capacity, and ROI, following Tang Mu's framework in woshipm.
XBOT frames the alternative around a task-defined product: a coffee robot operating as a small independent coffee shop at a commercial location. Tang Mu says XBOT made more than 500 coffee robots last year, has deployed more than 1,000 globally, and has cumulatively made about 4 million cups. Real orders bring customers, payment pressure, equipment failures, replenishment demands, output-stability requirements, and single-location ROI into the same test. Its Robot as a Service model also depends on continuing operations and revenue from connected supplies, coffee beans, and the supply chain. Wider deployment should wait until that operating loop holds.
Treat weight, safety, and automation as release gates rather than demo details. Guo Renjie said his team cut a robot from 15 kilograms to 8 kilograms through in-house joint modules; that kind of change should be checked against field handling, failure modes, and the tasks the machine can complete without intervention.
Write down the unanswered questions before scaling.
Guo Renjie recorded more than 300 questions after discussions with 40 employees. Product teams can use the same discipline: define human handoff rules, test safety around people and animals, and measure where autonomy stops. His early skeleton tests with cats show why edge cases belong in validation, not post-launch support. For home deployment, his stated prerequisites are safety and full automation. Do not infer market readiness from mobility or multimodal interaction alone: he calls those the only two mature embodied-AI capabilities, while no single 0-to-1 product has exceeded 5,000 unit sales.
For readers outside China
- Availability: The source material does not provide a general international purchasing guide for the robot platforms discussed. XBOT says it has deployed more than 1,000 coffee robots globally, but the countries, sales channels, and availability to individual buyers are not disclosed in sources. Zhiyuan Robot says Genie G2 robots entered Longqi Technology's mass-production factory in Nanchang, Jiangxi; this describes a deployment rather than a retail offering.
- Pricing: No prices, subscription rates, leasing terms, or Robot as a Service charges are disclosed in sources. XBOT describes its Robot as a Service model as continuing operations and revenue from connected supplies, coffee beans, and the supply chain, but does not disclose commercial terms.
- Closest Western equivalents: The source material does not identify or compare these products with Western equivalents.
- Data residency: The source material does not cover data residency, cloud hosting locations, cross-border transfer, retention policies, or enterprise data controls. It does say that XBOT uses AI employees to analyze operating data, location data, order heat maps, and operating status in group chats, but where that information is stored and processed is not disclosed in sources.
Sources
- ifanr 机器人运动会散场后,自变量给机器人建了一座「虚拟训练场」 https://ifanr.com/1677048
- woshipm 对话郭人杰:在绝对0-1的赛道里,家庭具身智能没有大哥 https://woshipm.com/share/6433356.html
- ifanr 智元机器人「奥运」首秀封神,戴着工牌狂揽 18 枚金牌 https://ifanr.com/1676970
- woshipm 如何设计自然的具身人机交互 https://woshipm.com/embodied/6457610.html
- woshipm 从人形概念到咖啡机器人,具身智能如何跨过商业化鸿沟 https://woshipm.com/it/6453574.html
- geekpark 「AstraTennis」时刻:机器人在全球直播中打了一场真正的网球 https://geekpark.net/news/369283
The evidence: 40 facts from 5 Chinese articles
Each line below was extracted from the article it sits under, in Chinese, before any of this was written. The writing is done from these and never from the source prose - that separation is structural, not a promise. How we work.
geekpark「AstraTennis」时刻:机器人在全球直播中打了一场真正的网球
- On August 22, a humanoid robot played a tennis exhibition match at the opening ceremony of the second World Humanoid Robot Games.
- The robot was developed by Galbot.
ifanr机器人运动会散场后,自变量给机器人建了一座「虚拟训练场」
- X-Square Robot released the autoregressive world model WALL-SS as humanoid robots move from competitions toward real-world applications.
- WALL-SS generates a low-resolution preview of future scenes before progressively adding details such as object contours, gripper opening and closing, and contact positions.
- In an action-sensitivity evaluation, WALL-SS scored 0.290, Cosmos3 scored 0.044, and all other models scored 0.
- The action-sensitivity metric measures a model's sensitivity to changes in actions.
- WALL-SS scored 0.539 for trajectory accuracy, the highest score among all compared models.
- WALL-SS's visual backbone, InfinityStar, scored 0.251 for trajectory accuracy.
- The WALL-SS research team retained training data covering missed grasps, slips, collisions, re-grasps, manual takeovers, and failure recovery.
- During training, WALL-SS adds perturbations to historical information and requires the model to predict the correct future from error-containing inputs through a method called scale-by-scale dream forcing.
- In a 60-second water-pouring simulation, WALL-SS kept frame-by-frame robotic-arm trajectory error below 0.5% of the image diagonal.
- Online visual-dynamics policy alignment increased WALL-SS action following from 0.264 to 0.290 and trajectory accuracy from 0.512 to 0.539.
- Online visual-dynamics policy alignment reduced WALL-SS cross-segment boundary error from 0.118 to 0.104.
- The research team tested five WALL-WM policy versions across six tasks and 20 matched initial-state groups in WALL-SS and on real robots, producing 600 pairs of closed-loop results.
- Of the 600 virtual and real-robot result pairs, 527 had matching final success or failure outcomes.
- For 30 task-and-policy-version combinations, virtual and real success rates had a correlation coefficient of 0.926 and a mean absolute error of 0.062.
- In real-robot tabletop dual-arm tasks, WALL-SS's action expert achieved an average task-progress score of 69.1, compared with 49.6 for π 0.5, 44.1 for DreamZero, and 34.0 for LingBot-VA.
ifanr智元机器人「奥运」首秀封神,戴着工牌狂揽 18 枚金牌
- Zhiyuan Robot led the medal table at the World Humanoid Robot Games with 18 gold medals, 16 silver medals, and 12 bronze medals.
- Zhiyuan Robot's dexterous hand won 7 gold medals across eight specialized events.
- Zhiyuan Robot won 6 of the 12 gold medals in task events and 5 gold medals in competitive events.
- Zhiyuan Robot competed for the first time at the World Humanoid Robot Games, using mass-produced models rather than models customized for the competition.
- Zhiyuan Robot's Expedition A3 performed with Olympic champion Ding Ning at the opening ceremony.
- Expedition A3 won Zhiyuan Robot's third gold medal at the event in the martial arts Tai Chi competition.
- Zhiyuan Robot's Lingxi X2 won the gold medal in the 100-meter obstacle race.
- The fire emergency scenario event required robots to identify hazardous materials, shut down abnormal valves, and grasp and use a fire extinguisher.
- The fire extinguishers used in the fire emergency scenario event weighed 4.5 to 5 kilograms.
- Zhiyuan Robot's library-event team, formed with Tsinghua University and Shanghai Jiao Tong University, was required to complete outbound transport, book shelving, and identification and correction of misplaced books.
- Zhiyuan Robot won first place in total score in the 2026 WorldArena Track 1 world-model perception and action-response competition with its GE-2 world model.
- In April, Zhiyuan Robot presented seven deployment-oriented productivity solutions for production-line loading and unloading, industrial handling, logistics sorting, guided tours and shopping assistance, service retail stations, security patrols, and industrial and commercial cleaning.
- Zhiyuan Robot founder Deng Taihua defined an XYZ curve for embodied-intelligence development: X spans 2022-2026, Y spans 2026-2030, and Z begins after 2030.
woshipm从人形概念到咖啡机器人,具身智能如何跨过商业化鸿沟
- Tang Mu is the founder and CEO of XBOT.
- At the 2026 AI Product Conference, Tang Mu discussed how product managers can enter the embodied intelligence field.
- Tang Mu studied applied mathematics at university and became a designer after graduating.
- Tang Mu worked at Kingsoft for two years before joining Tencent, when Tencent had fewer than 200 employees.
- Tang Mu spent 10 years at Tencent and helped introduce interaction design and user research roles and methods into internet product teams.
- Tang Mu left Tencent in 2013 and joined Xiaomi to work on smart hardware.
- At Xiaomi, Tang Mu worked on Xiaomi Router, Xiaoai Speaker, and Xiaomi VR all-in-one headset products.
- Tang Mu later started a business and entered the robotics sector.
woshipm对话郭人杰:在绝对0-1的赛道里,家庭具身智能没有大哥
- Guo Renjie is the founder and CEO of Lexiang Technology.
- Lexiang Technology's brand is Zeroth Yuandian.