Multimodal AI Is Not Just an Upgrade to Coding, But a Dual Reconstruction of Compute and Algorithms
Community Discussion · Tracks

Multimodal AI Is Not Just an Upgrade to Coding, But a Dual Reconstruction of Compute and Algorithms

Tian JiTian JiJul 202026/07/19 75 views

I basically agree with Lin Dahua's judgment at WAIC, but I need to add a key perspective: The Coding field has entered the "engineering optimization" phase, while multimodal is still in the "brute-force exploration" phase. Although both are core battlefields for large model implementation, the maturity gap in their tech stacks is at least one generation.

I've recently been running a multimodal RAG (Retrieval-Augmented Generation) project. In practice, I can achieve over 85% task completion rates with pure text Coding Agents, but once image and video inputs are added, performance drops directly below 50%. This isn't a problem of model capability, but rather that the underlying mechanism for modality alignment hasn't been truly solved.

The Difficulty of Multimodal Isn't "Understanding," but "Alignment"

Lin Dahua mentioned that multimodal is the next battlefield, which is technically accurate. But I want to break down why multimodal is harder than Coding from a practical level.

Coding is essentially structured text generation; the model only needs to understand the syntax tree and semantic logic of code. Multimodal requires the model to simultaneously understand pixel distributions in images, temporal signals in audio, and abstract concepts in text, establishing mapping relationships between these heterogeneous information types. The current mainstream approach uses CLIP or similar models as visual encoders, then aligns visual features to the language model's semantic space via Q-Former or linear projection layers. However, benchmarks show severe distortion in this alignment for fine-grained tasks.

For example, I recently ran comparison tests with GPT-4o and Claude 3.5 Sonnet: Given a complex scene image containing multiple objects, asking the model to describe "the color of the third object in the top-left corner." Both models exhibited positional offset errors and lacked precision in identifying object edges. This indicates that current multimodal alignment is essentially "coarse-grained semantic matching," not true pixel-level understanding.

In the Coding field, models can already form closed loops by verifying feedback through code execution. Multimodal lacks this natural self-supervised signal, and data annotation costs are much higher.

SenseTime's Technical Path: From "Full-Domain Perception" to "Dynamic Fusion"

In the interview, Lin Dahua emphasized that multimodal needs to solve the fusion problem of "perception" and "understanding." SenseTime's proposed solution, from a technical perspective, attempts to bypass the limitations of pure Transformer architectures by introducing spatial perception inductive biases.

I noticed their recently released InternVL 2.0 series adopts a dynamic resolution strategy for visual encoders, meaning it adaptively adjusts patch sizes for images of different dimensions, rather than uniformly resizing to a fixed dimension like traditional methods. This strategy shows significant effects in OCR and document understanding tasks. I previously tested their model and found it scored about 3 points higher than Qwen-VL-Max on DocVQA (Document Visual Question Answering).

But the problem is that this dynamic resolution strategy significantly increases computational overhead during inference. For edge-side deployment, real-time performance will be a bottleneck. Lin Dahua didn't discuss this directly, but I think if SenseTime wants to commercialize multimodal, they must make breakthroughs in model compression and quantization. Otherwise, they can only go the cloud-based large model route, where OpenAI and Anthropic already hold the first-mover advantage.

The Next Battle for Multimodal: From "Understanding" to "Generation"

Currently, mainstream multimodal models focus on the "understanding" side, i.e., taking image/video input and outputting text descriptions or answers. But the real battlefield is "multimodal generation," meaning generating images, videos, audio, or even cross-modal editing based on text instructions.

Tools like Cursor and Copilot have emerged in the Coding field, allowing developers to see real-time feedback while writing code. If multimodal generation can also achieve "real-time interactive editing," such as modifying an object in an image while describing it, that would be a truly disruptive application. But current diffusion models and autoregressive models lag far behind in generation efficiency, especially for video generation, where single-frame inference time is still in seconds, let alone real-time.

Regarding Lin Dahua's statement that "multimodal is the next battlefield after Coding," judging by the pace of technological evolution, I believe it will take at least 2-3 years to reach the maturity level of the Coding field. But there is one thing we can do now: Pay attention to data construction methods for multimodal alignment, especially using synthetic data for fine-grained annotation. This is currently the lowest-cost, most certain breakthrough.

[!tip] Action Advice

If you want to enter the multimodal track now, don't start directly with model training. Start with data engineering. Build an automated multimodal annotation pipeline, use existing visual detection models to generate pseudo-labels, and then perform precise annotation via human-machine collaboration. Once this workflow is established, regardless of which base model you use, you can quickly build advantages in specific domains. My team has validated this strategy in e-commerce scenarios, achieving results over 15% better than fine-tuning pre-trained models directly.

The future of multimodal belongs to those who can turn "modality alignment" from a black box into a white box. Don't be misled by current benchmark scores; the real tough battles are still ahead.

Original link: https://www.leiphone.com/category/yanxishe/uHHqpDWEcbFxyyvD.html

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts