
Before asking AI for math solutions, set up a verification process first
I spent the weekend messing around with using AI to solve math problems and verify calculations, stepping on quite a few pitfalls. It started when I saw a report titled roughly "AI is strong enough to chew through the hardest math problems," which also touched on risks. My first reaction was worry for myself. People doing chip verification fear tools giving beautiful results without an evidence chain underneath.
Let me explain two terms first. An AI model is a program that continues writing based on your input; you give it a problem, it gives a solution. A verification process is like chip testing, specifying what inputs are, what pass criteria are, and how failures are recorded. Without this process, asking AI math problems easily leads to a situation where it's confident, and I'm even more confident.
First, pick a short problem. Don't choose open-ended ones; choose ones with definite answers. For example, "What is the probability that 3 people randomly sitting in a row of 5 seats have no two adjacent?" Newbies tend to pick proving conjectures; these lack clear acceptance boundaries, and AI starts hallucinating.
Open your usual AI chat page, e.g., Claude web version, click New Chat. Don't ask for the answer yet. Input: Please restate this problem, listing knowns, goals, and boundary conditions. Expect to see it break the problem down into 5 positions, 3 people, non-adjacent, find probability.
Watch its restatement. This step is critical. If it understands "non-adjacent" as "can't sit together" but misses "any two people," correct it immediately: No two seat numbers can differ by 1. It's like reading a chip spec sheet; one wrong character throws everything off later.
Ask it to solve step-by-step. Input: Do not give the final answer directly. First list the counting method, then list the calculation formula. You'll see it write total cases and satisfying cases. If it skips steps, follow up with: Please explain why each step holds true.
Ask it to find counterexamples. A counterexample is a small example specifically designed to catch errors. Input: Provide a scenario easy to miscalculate and explain where the error lies. A good answer will mention whether seats are numbered or the permutation order of people. If it says there are no counterexamples, I generally don't believe it.
Verify locally. Open your computer terminal (the black box for entering commands), type python3, press Enter until you see >>>, then type import math and press Enter; then type math.comb(5,3) and press Enter to see a number. math.comb(5,3) means choosing 3 seats out of 5. If you don't have Python, use a phone calculator for intermediate numbers. The key is keeping your own evidence.
Archive. Create a new text file containing the original problem, AI restatement, AI solution, your verification, and final conclusion. Don't underestimate this step. In chips, if you don't archive, no one knows how it passed two weeks later.
I stepped on many pitfalls along the way. Recently, I used Harness for small workflow validation, understanding it as scaffolding to help models call tools. It can run tasks but can't judge right or wrong for you. In one problem, AI wrote the total cases as 5 choose 3, then multiplied by permutations of people, resulting in a probability greater than 1. Its tone was extremely steady. If I had only looked at the conclusion, I would have been led astray.
Another pitfall is self-reported status. It says "verified" but doesn't write how. In engineering, this is called a false pass. Power wall issues are the same; when compute can't scale further at a certain process node, the key becomes verification coverage and test stimuli. Math problems are similar; what really matters is the evidence chain.
This process has benefits and limits. Benefits: low barrier—if you can type, AI can help break down the problem; exposes vague expressions—if you can't clearly define "non-adjacent," it will misunderstand; suitable for practice—you can ask it what happens if conditions change, e.g., 6 seats instead of 5. Limits: it can't serve as authority directly; it might miscalculate combinations or cite non-existent theorems; long problems lead to missing conditions—as context grows, it forgets earlier boundaries; if you don't verify, it amplifies your false confidence.
My judgment is: treat AI as a problem decomposer and counterexample generator, not a math referee. Asking it to solve is fine; trusting its proof is not. At minimum, verify locally or recalculate key counts with code.
Reports mention that researchers at some AI company believe there is a >10% risk that needs attention. I won't comment on whether this ratio is accurate. But in engineering, I strongly agree with one point: the greater the risk, the more we need verifiable small processes, rather than emotional forwarding.
Next, try asking the same problem to two different models and compare where they diverge. Divergence points usually indicate unclear assumptions. Then, learn a bit of Python and script the combination counts. Let one answer leave evidence first.
📌 This article is compiled from Hacker News. Original: https://www.wsj.com/tech/ai/ai-math-millennium-prize-safety-openai-anthropic-05179825
Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.
Physix Frontier