Tops leaderboard with 85.21 score: Embodied LLMs evolve from 'understanding commands' to 'true comprehension'
Community Discussion · Policy

Tops leaderboard with 85.21 score: Embodied LLMs evolve from 'understanding commands' to 'true comprehension'

Yaoyao Product SelectionYaoyao Product SelectionJul 282026/07/28 74 views

Today I saw some news: Agibot's self-developed WITA-Omni Preview embodied omni-modal large model topped the DailyOmni leaderboard with a comprehensive score of 85.21, surpassing Qwen and Gemini. What does the number 85.21 mean for someone like me who deals with AI tools daily?

First, here's a data set: Last year, when I was doing cross-border e-commerce product selection, analyzing competitor videos using traditional methods meant processing a maximum of 20 videos a day. Later, using multimodal models for assistance, I could process 100. But the problem was that these models were mostly "single-modal"—looking at visuals was separate from listening to audio. When pieced together, situations often arose where "the visual was seen but the audio wasn't understood." WITA-Omni's "joint audio-video understanding" and "temporal reasoning" capabilities hit this pain point directly.

From "Locating Sound" to "Synchronized Audio-Visual Understanding": Where is the Technical Barrier?

As a self-taught tech enthusiast, I understand that the core difficulty for this type of model lies in "temporal alignment." For example, in a product demo video, a red button appears on screen while the audio says "this is the emergency stop key." Traditional models might separately recognize "red button" and "emergency stop key," but fail to precisely associate the two on the timeline. WITA-Omni's breakthrough is aligning audio and video signals in the time dimension, understanding "these things at this moment refer to this event."

Specifically regarding application scenarios, I thought of a few:

  • E-commerce Livestream Analysis: Simultaneously analyzing the host's script and the product display on screen to determine which scripts paired with which visuals yield higher conversion rates. Previously, this required recording screens, transcribing, and then correlating separately. Now it's done in one step.
  • User Feedback Video Processing: Unboxing videos shot by overseas users. Product defects appearing in the visuals and complaints in the audio can be automatically matched to generate structured reports.
  • Competitor Monitoring: Automatically scraping competitor videos, analyzing the match between their promotional focus and visual presentation, and identifying weaknesses in marketing strategies.

How Much Can Actual Conversion Rates Improve?

I did the math. Using this model for product selection, the team can compress 20 hours of manual analysis per week down to 3 hours, reducing labor costs by approximately 70%. But more critical is the ability to "discover blind spots." Previously, when we analyzed competitor videos for the Southeast Asian market using traditional methods, we never realized that local users emphasized "waterproof features" far beyond expectations. Later, using multimodal models to analyze subtitles, spoken words, and visuals in the videos, we discovered this signal. After adjusting our product selection strategy, the conversion rate for related categories increased by 15%.

However, there are still several practical issues with technology implementation:

  • Training Data Cost: Embodied models require large amounts of high-quality audio-video aligned data. Cross-border sellers can't do this themselves and must rely on platforms or third-party services.
  • Inference Efficiency: The leaderboard score is high, but it remains doubtful whether response speed in real-world scenarios can reach "second-level." Comparatively, running large models in the cloud versus local deployment offers vastly different experiences.
  • Adaptability: Terminology and scenarios vary greatly across industries, requiring fine-tuning. For instance, industrial goods and consumer goods have completely different video content structures; applying general models directly may result in discounted performance.

Trend Prediction: Over the Next Year, Embodied Large Models Will Reshape "Data Analysis" in Cross-Border E-commerce

My judgment is that within the next 12 months, capabilities like "joint audio-video understanding" will move from labs to commercial applications, exploding first in cross-border e-commerce, content moderation, and intelligent customer service.

The reason is simple: Competition in cross-border e-commerce has shifted from "price wars" to "information asymmetry wars." Whoever can extract user needs, competitor weaknesses, and marketing trends from massive amounts of video faster and more accurately will gain the upper hand. Models like WITA-Omni essentially solve the "information asymmetry" problem—turning videos that previously required repeated human viewing, recording, and comparison into structured data matrices.

But note that technology itself is just a tool; the true value lies in "how it is used." For example, combining temporal reasoning capabilities, one could create a "video content timeline," automatically annotating key information at each time point to help operations staff locate details quickly. This productization capability is more important than purely chasing leaderboard scores.

[!note] From a practical operational perspective, leaderboard scores can be understood as "theoretical peaks," but landing effects depend on engineering capabilities and scenario adaptation. An 85.21 score indicates that basic capabilities are sufficient; the next step is seeing who can turn the model into a useful tool.

Finally, I'd like to say to my peers: Don't just stare at the leaderboard. Think more about which parts of your business scenarios are troubled by "humans repeatedly watching videos." As long as you can solve this problem, whether it's an 85-point or 80-point model, it can save you real money.

Over the next year, I predict a batch of AI focused on "video understanding + e-commerce operations" will emerge...

Original Link: https://www.ithome.com/0/982/739.htm

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts