Physix Frontier · News Briefing Card (Leiphone · Aug 28, 2026)

Daxiao and HKU Unveil StreamPI for VLA Temporal Understanding

KEY FACTS

  • Daxiao Robotics and the University of Hong Kong introduced StreamPI to address single-frame decision flaws in Vision-Language-Action (VLA) models.
  • The system employs streaming caching and random-interval training to achieve temporal modeling without increasing parameter counts.
  • In real-world tests, the success rate for the cup-guessing task improved from 46.7% to 80.0%.
  • Success rates on long-horizon LIBERO-Long tasks rose from 92.4% to 95.0%.

KEY DATA

46.7%→80.0%Cup Guessing Success Rate Improvement
26.7%→63.3%Rolling Object Grasping Improvement
60.0%→92.0%Paper Cup Insertion Success Rate
92.4%→95.0%LIBERO-Long Success Rate

PHYSIX OBSERVATION

StreamPI marks a paradigm shift in embodied intelligence from static perception to dynamic spatiotemporal understanding. Its lightweight architecture proves that the temporal dimension can be augmented without scaling compute power, directly addressing robots' lack of 'process memory.' This capability to model the continuity of the physical world is a critical step toward breaking current VLA generalization bottlenecks and achieving true autonomous manipulation.

Source: Leiphone report