Agent Software Pilots Complete: Don't Just Chase the Hype
Spent the weekend tinkering with a pilot of agentic software for a client. Stepped on quite a few landmines.
Conclusion first: It depends. It's suitable for teams that already have system interfaces for tickets, knowledge bases, CRM, etc., and are willing to invest people in defining boundaries and logging. It's not suitable for small teams expecting an upgrade just by slapping on a chat box. The role of agentic software is to break a task into steps and then call systems to do the work; it's not just chatting. MCP is one protocol I used this time to connect external systems; you can roughly understand it as a standard for power outlets.
This time, I mainly looked at tools, processes, and governance. The tool layer involves models and interfaces; the process layer checks if business steps can be completed end-to-end; the governance layer deals with who approves, how to trace, and whether errors can be rolled back. The core conflict is that business wants speed, IT wants stability, and frontline staff want to avoid taking the blame.
This week, the Ministry of Industry and Information Technology (MIIT) released the "Implementation Plan for the 'AI + Software' Special Action," proposing that key software will fully achieve intelligent upgrades by 2030. I treated this document as a mission brief and stress-tested it against a client's internal after-sales ticket assignment process.
On Saturday morning, I built an agent in WorkBuddy using drag-and-drop nodes. I set up ticket description reading, connected MiniMax for intent recognition and priority judgment, and then generated draft replies for frontline engineers. On the first run, intents were recognized, but it confused "system login failure" with "account locked," and priorities were a mess.
I thought the small model was inadequate, so I switched to the test environment, but the issues became more obvious. The ticket system requires simultaneous updates for status, owner, and SLA; missing one causes an error. Here, SaaS refers to cloud-based software services, not local .exe files. In the afternoon, I got stuck on context. The test environment could only simulate data, and many historical rules weren't synced. My testing showed that the most time-consuming part was linking "this issue" with "the previous issue." The model only saw the current ticket and didn't know about past handling records. Later, I created a context summary table, bringing in the status, owner, and resolution of the last three tickets, which finally made the node turn green.
There's a "Run History" feature in the interface where you can see the input/output of each node. This feature saved me. Previously, when the model misjudged priority, I had to guess. Later, I found that the old ticket summary wasn't included in the context; the model only saw the current sentence and didn't know what "still the same problem as last time" referred to. After cleaning this up, the judgment stabilized.
Anomaly detection surprised me a bit. I threw in two test tickets with similar descriptions and the same customer. The agent flagged them as "suspected duplicate tickets for the same issue" and cited the original ticket numbers. This isn't super advanced, but it's useful for initial frontline review. It can block things that clearly shouldn't proceed to the next step, but it can't replace humans.
On Sunday, I supplemented logs according to the SB 53 checklist. SB 53 is California's AI Safety Act. I recently started breaking it down into a checklist to use as an external template. After completing it, every step records the model version, prompt, which system was called, and whether human confirmation occurred. It looks like a maturity questionnaire, but it can save your skin when things actually go wrong.
I also tried integrating AI coding assistants, letting developers generate some validation code calls. It did generate code, but it didn't meet the client's coding standards, exception handling, or log formats. Fixing it was slower than writing it by hand. This shows that AI coding is good for scaffolding, not for direct delivery.
Xinhua News Agency mentioned that by 2028, adoption will cover 20,000 large-scale software enterprises above designated size, with a cumulative implementation of 100 intelligent technical transformation projects for software enterprises.
This statement feels distant from ordinary users. On the ground at the client site, it means fewer ticket transfers, shorter reply drafts, and less page flipping during upgrades. The plan emphasizes "application-driven, innovation-led, secure and controllable, ecosystem collaboration." I care more about whether "secure and controllable" and "ecosystem collaboration" can translate into logs and rollbacks.
The benefits are direct: When structured data is clean, generating ticket summaries and meeting notes really saves time; visualized processes allow business and IT to argue over the same diagram; anomaly detection catches low-level errors. The downsides are hard: Dirty data means you can't skip any cleaning effort; access boundaries, error handling, and rollback mechanisms require manual configuration; model outputs still need review and cannot be taken directly as conclusions.
My takeaway from this pilot is: Don't just look at the hype. For agentic software to land, the key is setting up boundaries, logs, and rollbacks. Only when these are in place can we talk about engineering-grade solutions.
Physix Frontier