Turn user feedback into AI prototypes with this workflow
Community Discussion · Tracks

Turn user feedback into AI prototypes with this workflow

YimingYimingSep 122026/09/12 71 views

Last Wednesday, sales forwarded me two hundred tickets asking if we could sort them by whether they need to be scheduled for next week. Doing it manually takes two hours and is prone to errors. I lead a team of thirty; my biggest fear is treating AI as a toy. This direction is worth pursuing, but it must integrate into our workflow. I built a feedback triage prototype using Claude. The process isn't complex—you can copy it.

The task is to turn user feedback into machine-readable output that can be checked, then feed it into the product board. The prototype's goal is first to get the judgment running and allow for human review.

Start by getting a minimal flow working.

Pick a small entry point. Don't tackle big topics like "Using AI for customer insights." Instead, change it to categorizing a week's feedback into five types: clear requirements, experience complaints, configuration inquiries, suspected bugs, and irrelevant for now. Categories must be mutually exclusive; any overlap confuses the model.

Prepare 30 sample entries. Export tickets to CSV. A CSV is just a comma-separated table that Excel can open. Fields only need ID, original text, and source. Delete customer names, phone numbers, and contract amounts. If privacy isn't handled, legal will block you.

Define success criteria. Don't write "looks accurate." Set a standard: manually check ten samples, at least nine directions must be correct; no suspected bugs should be missed; each entry needs a reason and suggested action. The reason explains why the model categorized it this way, and the suggested action is what to do next.

Open the Claude chat window and start a new session. Paste the following content into the input box.

You are a B2B product analysis assistant. Please classify each feedback item into one of the five categories above. Output JSON with fields: id, category, reason, action. Judge based solely on the original text; do not fabricate information.

Then paste the samples. After sending, you should see a set of entries enclosed in curly braces. Expect each to have an ID, category, reason, and action. If the format is messy, add a line saying: "Please strictly output a JSON array, no extra explanations."

Manually review ten entries. Add three columns in Excel: Model Judgment, Human Judgment, Consistent? If inconsistent, don't blame the model immediately. First check if the category definitions are unclear or if the original text lacks context.

Integrate into scheduling. Drop high-priority actions into the board. For example, route suspected bugs to technical support, and put clear requirements into product reviews. Without this step, AI output is just pretty text.

Solidify into a weekly report. Feed new feedback in the same format, letting the model output summaries and classification tables. Run it manually first; automate once stable. I've been trying AI Agents recently—programs that automatically pull data, run models, and write reports. Beginners shouldn't jump straight to Agents; first, master running those 30 samples.

Newbies often fall into these traps. Too few samples—treating three or five as gospel truth regarding boundaries. Categories too similar—for instance, mixing experience complaints with clear requirements makes prioritization impossible. Only looking at demos without considering production feasibility. Customers might want automatic sunroof adjustments; AI can summarize the requirement, but it won't make decisions on structure, cost, or supply chain for you. Also, dirty data—wrong date formats, duplicate submissions, screenshots not converted to text. Cleaning data takes more time than modeling.

The pros are direct. Fast—you can get a flow working in an afternoon. Checkable—you can use manual review results as proof. Integratable into existing workflows. The cons are also obvious. Models cannot replace product judgment, nor guarantee every entry is correct. Sensitive materials shouldn't be casually thrown into the cloud. Outputs require human spot-checking. I've used Claude for about a month. My feeling is that processing structured feedback is indeed convenient, provided the feedback itself is regular.

In Andrew Ng's AI Engineering Skills Map, there's a judgment I agree with: AI engineering requires repeatedly building, checking, and deciding the next step. Reportedly, this skills map is based on analyzing over ten thousand job postings and interviewing many experts. This work must land on integrating every piece of feedback into scheduling.

If stable classification outputs are achieved, try two things next. Add a confidence field. Confidence is how sure the model is about its judgment. Low-confidence items go to humans, reducing false positives. Move the weekly report from the chat window to a spreadsheet for sales and customer service. Only after that is it worth considering cloud GPUs, internal models, and Agent orchestration.

Over the next six months, being able to integrate model outputs into team workflows is more valuable than just writing prompts.


📌 This article is compiled from Hacker News. Original: https://twitter.com/andrewyng/status/2098459474608672916

Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.

2 replies

?
Ctrl + Enter to reply
Ling Xi
Ling XiSep 12

Tried it out. The scariest part is unstructured feedback. Don't rush into AI; nail down the dirty work of cleaning and labeling first to keep things stable.

Factor Miner
Reply to Ling Xi

User feedback is too noisy. With insufficient sample size, you simply can't train stable factors. Don't get fooled by the prototype.