Agents Write Code, But Who Tests on Mobile?
Community Discussion · Tracks

Agents Write Code, But Who Tests on Mobile?

LuguoLuguoSep 112026/09/11 88 views

Today I saw Alibaba Qoder's Mobile Use plugin, and my first reaction was: finally, someone is addressing the last mile of mobile App development. Many Coding Agent promotions claim they can modify Android, HarmonyOS, and iOS code, which sounds powerful, but anyone actually building Apps knows that finishing the code is just the beginning. Did the build pass? What does the page look like? Is the button the user mentioned the same control the agent clicked? Did clicking it lead to the wrong place? If these aren't solved, the Agent is still stuck at the "help me write a snippet" stage.

I've been testing Claude and Perplexity lately, using them for some small frontend tasks. Text-layer stuff is indeed fast, but it breaks down once you hit the actual device interface. The reason isn't complex: what the model sees differs greatly from what the user sees. In the code repo, it's a function; on the phone, it's status bars, pop-ups, permission prompts, list refreshes, and animation delays. If a mobile Agent can only read diffs but not runtime states, it tends to write patches that are syntactically correct but experientially wrong.

Looking at recent AI Agent funding and new products, it's easy to mistake demos for products. For the mobile line, the real difficulty lies in connecting devices, OS versions, permission dialogs, widget trees, logs, and regression test cases. Only when an Agent can prove its changes work does it resemble solving an engineering problem.

So my judgment on this type of plugin is: the key is whether it can solidify the verification process. Supporting Android, iOS, and HarmonyOS is just the entry ticket. We need to see if it can stably take screenshots, grab logs, identify widgets, simulate clicks, and feed results back to the model. Especially for HarmonyOS, ecosystem adaptation and debugging entry points differ from Android/iOS. If it's just a conceptual demo, it's meaningless; if it can connect to real devices or cloud real devices, then it has some merit.

I suggest teams don't rush to let Agents auto-merge code. Treat it as an automated acceptance tester instead: read-only interfaces, run fixed paths, generate evidence packages. Give it lower modification permissions, and submit PRs only after human confirmation. Wait until it can smoothly run through the flow of "no crashes after changes," "clicks go to the right place," and "screenshots match," before talking about auto-merging.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts