Easy to Do Backflips on Stage, Hard to Enter Homes
I've been pondering this question for a while: Why do robots that flip, dance, and serve tea at exhibitions seem quite capable, yet when it comes to buying one for home use, everyone's first reaction is to shake their heads?
Over the past eighteen months, funding for humanoid robots has multiplied several times. I don't doubt this figure. But behind the capital and exhibition hype, what's truly worth digging into is another thing: Is the gap between "performing an action at an exhibition" and "accompanying someone at home all day" really just a few generations of technical iteration away? Or are they fundamentally two different paths, just lumped into the same basket?
From my risk control perspective, the core contradiction here is almost identical to deploying an anti-fraud model. Exhibitions are structured environments: fixed lighting, flat floors, audiences separated by fences. Robots only need to execute scripted action sequences, essentially running on a "controlled condition" test set. Home environments, however, are completely open unstructured spaces. Lighting changes, carpet friction coefficients vary, kids' toys are scattered on the floor, and worse, everyone's behavior doesn't fit any script preset. It's like offline backtests looking beautiful until real traffic hits, revealing flaws as feature distributions drift, false positive rates spike, and black market actors never play by the test set rules.
Zhang Fumin's statement was quite accurate: companion robots need to simultaneously tackle four technical pillars: motion, interaction, memory, and touch. But my first reaction after reading it was that the factor determining whether they can enter homes might not be "motion," as most assume. Making joint motors smoother is an engineering problem solvable with money and time. The real hurdle lies in memory and interaction; these determine whether a robot seems "human-like" or "machine-like."
Interaction requires locally deployed large models for low-latency multi-turn dialogue, which involves a computational deadlock: Large models require training clusters of tens of thousands of GPUs, but robot bodies can only run embedded platforms like Jetson. Bridging this gap requires a full chain of model compression including pruning, quantization, and distillation. I encounter similar issues in anomaly detection: models that run fine on servers either lose unacceptable accuracy or suffer latency delays causing users to uninstall when moved to edge devices. Selling this generation of robots cheaply, resulting in dialogue responses lagging by three seconds, is more off-putting than stiff movements.
I think memory is the most underestimated aspect. Long-term memory plus emotional state machines, retaining personalized information across sessions—this sounds like just adding a database to an LLM, but the implementation scenarios are full of pitfalls. In risk control, user profiling dimensions rarely exceed dozens or hundreds of features, mostly numerical. Human conversational memory is semantic-level: yesterday's joke, a colleague's name mentioned last week, tone changes when angry. Maintaining consistency across sessions places extremely high demands on state management. In my tests, slightly longer contexts lead to information loss. Extracting task info from chat flows is most feared due to context breaks. This problem will only be more severe in robots because their input isn't just text, but multimodal fusion of vision, touch, and environmental state.
Then there's the simulation-to-reality gap, an old issue I mentioned in last week's medical robot article. Domain randomization in simulations stirs up friction, lighting, and sensor noise, essentially trying to exhaustively cover real-world uncertainty. But real-world uncertainty evolves like black market tactics; there's always a path you haven't enumerated. Accumulating years of human interaction experience in hours via simulation sounds great, but the distribution of experience differs by orders of magnitude from reality. Knocking over a cup, a cat jumping on a table, an elderly person suddenly pausing—these extreme long-tail events cannot be covered even with a hundred random seeds in simulation.
Also, the fragmentation of data itself. Two-finger, three-finger, and five-finger dexterous hands each have their own datasets. Data collected from different embodiments cannot be reused. To me, this is typical reinventing the wheel. In anomaly detection, each firm has its own feature engineering tricks, but at least the data schema is unified, allowing different algorithms to run on the same data. Robotics is different; embodiment form locks the data format. Switching hardware makes open-source datasets unreproducible, meaning specimens are collected but cannot be shared. Unless this landscape changes, industry-level general datasets won't emerge, and firms will remain stuck in internal competition within their small circles.
Therefore, I tend to judge that industrial scenarios will succeed first. Warehousing, inspections, and production lines are semi-structured, with predictable rules and clear task boundaries. Even if problems arise, humans can easily intervene. Why are home scenarios so difficult? Because the definition of the scenario itself is "no standard scenario." For industrial robots, you can define accuracy, success rate, and cycle time. For companion robots, you must answer: What is "normal companionship"? This definition varies across families, cultures, and age groups. The most profound lesson in my risk control career is that identifying anomalies is easy, but defining "normal" is extremely hard, because normality itself is a floating horizon.
Flipping at exhibitions proves the upper limit of engineering. Entering homes bets on the understanding of "normal." The latter is the true high wall.
Physix Frontier