Community Discussion · Policy

Embodied ICL Trend: Don't Treat Context as a Silver Bullet Yet

ZhulongZhulongSep 62026/09/06 95 views

Last Wednesday, I was standing by a test bench at a friend's company, watching a robotic arm organize a batch of charging gun heads with various shapes. It didn't use overly complex fixtures beforehand; instead, there were just a few reference photos and a task description nearby. The setup was changed on-site to let the model look at a few examples before acting. My friend said that in the embodied AI startup circle, this approach is called ICL, or In-Context Learning. I wasn't particularly excited after hearing it; my first thought went to my recent experiences using Robotaxi and FSD.

In 2020, GPT-3 shook NLP by learning new tasks from just a few examples. Six years later, the same narrative has been transplanted onto robotics. The startup mentioned in this Quantum Bit article is Cocoa Matrix, founded in April 2026. Founder Gao Yuxiang mentioned that when they discussed starting up in January of this year, the team had already decided to bet on ICL. This timing is interesting. The team treats context as a new scaling direction while hardware bodies are still competing fiercely.

From an implementation perspective, this is actually a shift from "training a larger model" to "organizing more effective on-site memory." Context in language models consists of prompts, examples, and historical dialogue. Context for embodied models is much more complex. Vision, torque, tactile feedback, proprioceptive state, task goals, environmental constraints, and records of previous failures can all enter the context. It's far more complicated than just stuffing in a few images and writing "do this."

I've recently used WorkBuddy for process organization and gained an intuitive understanding. Throwing multi-source inputs into the model all at once often results in unstable output. Breaking it down, processing step-by-step, and then merging yields much more reliable results. Embodied ICL cannot avoid this engineering problem either. When a robot makes a mistake, the cost is higher than when text is wrong. Text can be regenerated, but a robotic arm might crash, a gripper might damage something, or a person could get injured. These past few days, I've also been testing general-purpose large models for task decomposition. They list steps efficiently, but start drifting when physical constraints come into play. For example, if asked to plan an assembly action, it can write out what to pick up first and where to place it last, but it doesn't necessarily know where torque, tolerances, or gripping sequences will jam up on a real bench.

On the autonomous driving front, I've tested Robotaxi, FSD, Waymo, and Pony.ai over the past three weeks, and the experience feels increasingly similar. What often widens the gap is whether long-tail scenarios have been remembered, whether rules have been constrained, and whether fallbacks have been recorded, replayed, and attributed. Urban NOA is the same. Hardware is becoming easier to buy—LiDAR, 800V systems, battery packs, sensors—the spec sheets look good, but that doesn't equal stable experience. Real-world data shows that the same sensor solution can result in user perceptions differing by a generation depending on different software versions and operational boundaries.

Embodied intelligence is now going through the same phase. Capital is starting to look at hardware, engineering systems, and capabilities for data recording, replay, and attribution. Evaluations are cooling down too. That previous article about the embodied AI "Gaokao" (college entrance exam) mentioned that humans score 100 points, while the strongest model only scores 12.8. This conclusion isn't pretty, but it has dampened the demo-driven atmosphere. A robot completing a few beautiful actions doesn't mean it can be repeatedly deployed.

If context is only used for showing off, its value is limited. A robot learning a new task by looking at a few examples on-site is indeed easy to spread. But once it enters factories, labs, logistics warehouses, or homes, problems become troublesome. Who provided the examples? Is there version control for the context? Which states did the model read at the time? Can errors be replayed? How is responsibility assigned? I usually pay attention to smart driving regulations, and my first reaction to this kind of on-site learning capability is whether the traceability chain is sufficient. Autonomous driving accident reviews require checking data logs, model versions, perception results, decision paths, and remote takeovers. If embodied ICL wants to enter real production environments, it will eventually face the same issues.

Don't just build a bigger brain; also establish links for experience collection, replay, and attribution.

I agree with this statement. Liu Ziwei from Ropedia talks about Embodied Scaling Laws, focusing on human experience. This judgment is closer to reality than simply stacking models. An experience system needs to land on a full suite of tools for collection, cleaning, alignment, replay, evaluation, and failure attribution. Data volume is just the entry point. If ICL is merely runtime retrieval, lacking recording, replay, and attribution, it will ultimately become just advanced prompting.

So my attitude toward embodied ICL is: not opposed, but don't deify it. It has the potential to become an important engineering path for embodied intelligence because robots naturally need to utilize on-site context. However, it won't replace training, simulation, real-machine data, and engineering constraints. Its more likely position is to string these things together, allowing robots to retrain less and adapt faster in limited scenarios.

Vertical scenarios will benefit first. Fixed workstations, standard bins, lab sample processing, and warehouse replenishment—these tasks have clear boundaries, definable contexts, and recoverable failures. Home scenarios are the most hyped but the hardest. Every home is different, items are different, habits are different, and safety requirements are high. So-called general-purpose robots, in the short term, look more like vertical experience compounding; one brain dominating all sites is still early days.

From an implementation perspective, the value of context lies in becoming auditable task state. If a piece of context only makes the model smarter in videos but cannot record, cannot replay, and cannot explain itself when things go wrong, it is still one layer away from engineering. Whether the embodied ICL track can succeed depends on this.

Looking ahead, I judge that in the next one to two years, the watershed for embodied startups will fall on stable context engineering capabilities: accurate collection, clear organization, stable evaluation, and investigable incidents. By around 2027, players who secure orders will likely be companies that turn ICL into deployable, operable, and compliant task memory systems. Just pitching ICL as a general-purpose brain isn't enough.

1 replies

?
Ctrl + Enter to reply
Zhi Wei
Zhi WeiSep 6

Once real-device latency increases, the visual feedback in the context completely desyncs. I've fallen into this trap before.

Embodied ICL Trend: Don't Treat Context as a Silver Bullet Yet - Physix Frontier Forum