Physix Frontier · News Briefing Card (Leiphone · Aug 21, 2026)

Xingdong Era: World Models May Replace VLA in Embodied AI

KEY FACTS

  • Chen Jianyu proposes that world models could become the core paradigm for next-generation embodied intelligence, rather than serving merely as enhancement modules for Vision-Language-Action (VLA) systems.
  • He notes that VLAs rely on imitation learning, which limits generalization, whereas world models improve zero-shot task capabilities by understanding physical laws.
  • Xingdong Era unveiled the world's first world-action model integrating video and action prediction, launching approximately one year ahead of NVIDIA's solution.
  • The company's self-developed fully direct-drive dexterous hand achieves 10 clicks per second with a 24-kilogram payload, supporting complex fine manipulation and cross-body applications.

KEY DATA

10/secDexterous Hand Click Frequency
24 kgDexterous Hand Max Payload

PHYSIX OBSERVATION

Shifting from mimicking human actions to understanding physical laws is the critical leap for robots to break through generalization bottlenecks. World models enable robots to 'imagine' future scenarios, freeing them from reliance on seen data. This marks a transition in embodied intelligence from specialized automation to general cognition, making closed-loop evolution of hardware-software synergy the new competitive high ground for the industry.

Source: Leiphone report