
Make AI-generated numerical outputs auditable
Recently, I looked at a game project. The PM said equipment stats were AI-generated, tweaked by designers, and launched by operations. I asked why the stat was 38 and not 28. They said the model gave it. In that moment, I deducted points from the project. Anyone can tweak a model; how high is this technical moat? Not high. High is whether numbers can be traced, reproduced, and handed over to the next person.
On GitHub, there's a free manual called AI Workflow for Game Designers, subtitled No Fabricated Numbers, meaning don't invent numbers. The author is Minsoo Lee. It's a six-month field manual covering Claude Code, prompts, verification, and production memory. Simply put, Claude Code is an AI assistant that reads files, modifies files, and works according to rules within a project folder. Prompts are fixed instructions for AI; verification checks if numbers have basis; production memory keeps confirmed versions. I've been trying Claude Code these past few days and built a minimal workflow accordingly.
I first used web-based AI for comparison. I pasted a rule table and asked it to check for numerical issues. In about three minutes, it gave an analysis suggesting increases, decreases, and balance risks. It looked nice, but didn't tell me which document clause was the basis, nor which numbers needed human confirmation. This kind of output is just opinion, not an asset.
Second path: Put documents in the project folder and use Claude Code with fixed prompts to output a verification table. It took me about twenty minutes to set up, resulting in a CSV (an Excel-compatible table) with fields: Field, Value, Source File, Source Location, Verification Status, Pending Questions. In the first round, many source columns were empty. I tightened the prompt: if no explicit source, write 'unknown'. In the second round, all unsourced items were flagged. Slow, but the ledger emerged.
I built it along a minimal verification line. First, create a folder on your computer named game-validation. Open it and create three subfolders: docs, prompts, validation. Put product docs, stat tables, and meeting minutes in docs; AI instructions in prompts; AI outputs in validation. AI most often errs between "it thinks you know" and "you know it doesn't know."
Newbies shouldn't test on the whole project. Drag one equipment stat table into docs and add a page of rule explanations. Filenames should include version and module, e.g., equipment-v1.md. The cleaner the file, the easier the tracing later.
In prompts, create a text file validate-numbers.md. Hardcode the requirements: Answer ONLY based on docs. Every number MUST provide source file and location. If source not found, write 'unknown'. Do NOT fill in gaps or guess. Output a table at the end. The focus of this prompt is to prevent AI from being vague.
Open the terminal in the game-validation folder (the black window for commands). Start Claude Code. If the interface can read files, input the command: Read all files in docs and generate validation/numbers.csv according to prompts/validate-numbers.md. Expectation: It exposes uncertainties.
Open validation/numbers.csv. Focus on Source File and Source Location. Anything marked 'unknown' goes to designers for questions. Anything with sources gets spot-checked to see if the original text supports the number. Here, AI just structures ambiguous problems; responsibility remains with humans.
Write confirmed numbers into production-memory.md, including version, date, owner, and reason for change. Future changes append new records rather than overwriting old ones. This action seems clumsy, but the exit path is clear: who changed it, why, who approved, and how to rollback—all traceable.
The easiest pitfall is AI treating "based on context" as a source. Solution: Specify in the prompt that without precise filename and paragraph location, mark as 'unknown'. Another pitfall is messy file naming—today final, tomorrow final_new, day after tomorrow nobody trusts it. Solution: Categorize directories and version files. Yet another pitfall is teams skipping acceptance for convenience, which just swaps the black box from model to spreadsheet, meaningless.
A few days ago, I wrote about viewing AGI through cost accounting. Today's line applies that thinking to projects. My tests show this line saves limited time; the real value is turning AI output from "I feel" to "auditable." As an investor in hard tech and new energy, I care more about whether it enters cost accounting, including compute, manual verification, rework, and liability boundaries. Without these, no matter how lively the business model, it's hollow. How high is this technical moat? The model itself isn't high; process discipline is. Exit paths are clear: short-term internal efficiency tool, long-term private audit platform.
Next, try swapping one stat table for two versions (old and current), let AI generate a diff explanation, and have owners sign off. Also, integrate verification status into weekly reports to see if rework frequency drops.
I will continue evaluating against this standard: only AI outputs that enter cost accounting are worth precipitating.
📌 This article is compiled from Hacker News. Original: https://github.com/eremes81/game-design-ai-practice-en
Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.
Physix Frontier