
Build a validation loop before buying AI compute
Last week at a product meeting, a colleague wanted AI to organize customer feedback and automatically remind sales. The boss's first question was how much compute power we needed to buy. I didn't engage. This direction is worth watching, but execution is key. After three years of entrepreneurship, what scares me most is talking about scale right off the bat. First prove it saves people, then talk about cloud bills.
If you can build a verification system and give proper direction, AI will go where the oracle is cheap.
Here, 'oracle' isn't a crystal ball; it's "who scores the results." The cheaper it is to judge correctness, the more willing AI is to work.
So this tutorial does just one thing: building a verification loop from 0 to 1. Simply put, let AI do a small task, then use cheap rules or human review to judge if it's correct.
Option A is manual verification on a small sample. One model chat window, one spreadsheet, ten anonymized samples. Suitable for when you haven't decided whether to proceed. Option B is Cloud API verification. An API is the entry point for programs to call models, suitable for when you've confirmed a passing threshold and want to check latency and failure rates.
Look at external numbers structurally first. OpenAI invested about $300 billion in Oracle services, Oracle invested $55.7 billion in AI infrastructure, laid off about 21,000 people, with restructuring costs of $1.84 billion. This isn't telling you to spend that much, but a reminder: where money goes, the organization follows.
Let's run through it specifically, using "organizing customer complaints" as an example.
1. Choose a narrow scenario. Don't do "AI improves efficiency," only do "extract problem type, sentiment, and suggested action from feedback." The narrower, the easier to verify.
2. Prepare 10 real samples, which must be anonymized, meaning privacy info is removed. Delete all names, phone numbers, companies, and order IDs. If you skip this step, everything after is a trap.
3. Build a scoring sheet. Columns: Original Content, AI Output, Is Correct, Can Go Live Directly, Time for Manual Fix, Risk. This sheet is the prototype of the oracle.
4. Open the model chat window, e.g., I used Claude for a month. Input a fixed prompt: "You are a feedback organizing assistant. Extract problem type, sentiment intensity, and suggested action from the content below. Output JSON only, no explanations." Paste the first item, hit send.
5. You'll see a JSON snippet in the interface, which is machine-readable structured text, like a small table. Expected structure looks like: {"problem":"Slow refund","sentiment":"High","action":"Escalate to CS Supervisor"}. If it starts writing prose, add: "Strictly output JSON."
6. After running 10 items, score them manually one by one. My habit: if fewer than 6 out of 10 are directly usable, don't do auto-reminders yet; go back and tweak prompts or change scenarios.
7. Try Option B next. Enter the cloud console (web backend), create a project, find Budget or Quota, set a daily cap. Use a program to call the model with the same batch of samples, recording time taken, failure counts, and output fields. Get the logs working first.
Hard budget caps are crucial. Newbies often think calling models costs pennies, but a bug in a loop can make the bill scary. I won't quote specific amounts, but I always set limits first.
Here are some real pitfalls.
- Just keywords won't work. I recently spent a week on regex and exclusion words before realizing negative phrases like "no refund" or "don't send coupons" cause false positives. Solution: Add exclusion words after keywords, then use regex to limit context.
- Without a scoring sheet, teams argue. Product says it looks good, Ops says it's too verbose, Boss asks if it can go live. Solution: Define the standard for "can go live directly" first.
- Treating the model as a decision-maker. Solution: Let it only draft; keep the send button for humans.
- Jumping straight to compute services like Lambda. I've been using Lambda for less than a week, but the conclusion is direct: without a verification loop, compute power only amplifies uncertainty.
After learning this, the next step is trying to semi-automate the "Is Correct" column in the scoring sheet. First filter obvious errors with regex and exclusion words, then have the model rerun, with humans only reviewing edge cases. You can also pipe results back into internal spreadsheets so sales sees less raw feedback.
📌 Compiled from Hacker News, original article: https://dilpreet.co/blog/2026/9/1/ai-goes-where-the-oracle-is-cheap
Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.
Physix Frontier