JD.com's Physical AI Data: Labs Shouldn't Copy It Blindly
Community Discussion · Policy

JD.com's Physical AI Data: Labs Shouldn't Copy It Blindly

Zhe Dan Bai DeZhe Dan Bai DeSep 92026/09/09 111 views

I compared JD.com's publicly available embodied data samples with lab protein purification operation logs and actually ran through them. Embodied data, simply put, refers to records left by robots or sensors after seeing, touching, and failing in real environments—not just text prompts. The core focus of JD's Physical AI Acceleration Plan is compute, data, and scenarios. It sounds like big words, but practically speaking, let's first see if the data format is usable.

During preparation, I only got public introductions and some samples; I didn't get the complete internal platform. I've been testing local AI inference nodes recently. Inference nodes are servers dedicated to running model predictions, separate from training large models. I tried splitting camera views, device states, and action instructions from the JoyInside AI Home and human-perspective datasets into three columns, then aligning them temporally. Action events in the samples have timestamps, but the coordinate systems are inconsistent with lab robotic arm logs, causing immediate sticking points.

Later, I converted wet lab operation logs into event streams, including opening lids, pipetting, centrifuging, and alarms. The model's performance on biological data is interesting: it looks at whether actions progressed from execution to completion. For example, if a pipetting action only has "execution" but no "completion," the model mistakenly assumes success. JD emphasizes scenarios, which aligns with my previous judgment on cold chain weather models: without hourly, micro-environment level data, scheduling automation is difficult. Lab scenarios are more fragmented; reflections on countertops, glove occlusion, liquid bubbles—these unstructured interferences eat up accuracy.

There were mainly two pitfalls. One is that the data feedback interface is unclear; I couldn't verify if it truly auto-labels failure samples into training sets. The other is that terminal connection descriptions are ecosystem-focused. The JoyInside plan connects over ten million terminals, but small teams care more about interface permissions, log fields, and privacy boundaries. The page only shows device online status and task completion, lacking failure reason fields. This state is dangerous in real-world operations. The after-sales network covering 100+ countries and 100,000 robot after-sales engineer positions over the next five years is a practical direction, because physical AI ultimately competes on on-site maintenance; launch PPTs don't solve problems.

Conclusion: It depends. If you're a large team with real equipment scenarios in logistics, home services, or robot after-sales, JD's physical AI route is worth watching, especially since data feedback and after-sales systems might fill industry gaps. If you're just a small lab, don't rush to integrate protein purification processes. The gap between dry experiments and wet experiments is still there. Stabilize local data field validation first, then talk about models.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts