Don't Pay Yet for Agents That Handle Tasks
As someone who habitually evaluates AI using reproducible workflows, I tried letting an agent handle tasks.
Han Xinyi said at the Bund Summit that the key shift in agentic commerce is moving from "capability demonstration" to "real transactions."
For this type of task, an agent can be understood as an AI that breaks down tasks. You give it a requirement, and it provides an answer, decomposes the task into steps, looks up information, calculates constraints, and finally delivers a confirmable draft. Referencing the ReAct paper's approach, let the model reason while calling tools. Tools include external capabilities like search, maps, and calculators. Our team budget is tight, so I care more about reproducibility.
I ran two approaches with the same requirement: a budget of 200 yuan for 5 people, finding nearby affordable and delicious restaurants.
Approach A was single-turn Q&A. Open a chat-capable AI, click the input box, enter the requirement exactly as is, and hit send. The expectation is that it will make recommendations. I tested 5 similar requirements; 3 only provided restaurant names and per-person costs, without distance, booking availability, or menu basis. It looked like an answer, but was actually hard to reproduce.
Approach B used three rounds of prompting, speaking to the AI in three steps. Round 1: Input a request to convert the requirement into a field table, including fields for budget, number of people, location, time, taste preferences, and mandatory exclusions. Observe the field table. Round 2: Input a request to list candidates based on the fields, requiring each candidate to cite a source; if no source, mark as uncertain. Round 3: Input a request to calculate total price by multiplying the number of people by the per-person cost, exclude any exceeding 200, and output a table with fields for restaurant name, per-person cost, estimated total price, distance, source, and items to confirm.
If the platform has Workflows, you can fix these three rounds. Click New Workflow, add an Input node for the requirement, a Search node to find restaurants, a Condition node for budget and distance, and an Output node requesting a table format. Workflows string these steps into a process, reducing free-form improvisation.
Testing the same 5 requirements, Approach B yielded verifiable fields for 4 cases, while 1 case could only provide candidates due to lack of real-time pricing. The difference mainly stems from experimental design, which broke down transaction intent into checkable variables.
Pitfalls lie in authorization and pricing. Don't click Confirm Payment or authorize orders immediately; first have it generate an order draft. Don't let the model report prices itself; prices, business hours, and inventory must come from tools or your screenshot confirmation. In prompts, write less "recommendations" and more "filtering conditions" and "output fields." If there's no search entry point, switch to low-risk scenarios, such as checking bus routes or buying coffee.
Following this logic, the next step is to try a real but small-value task: set budget, location, and time, have the agent provide three verifiable options, then manually confirm. First, ensure it leaves a chain of verifiable evidence.
Physix Frontier