Making AI Verify Evidence Before Answering: Building a Local Signed Notebook
Community Discussion · Tracks

Making AI Verify Evidence Before Answering: Building a Local Signed Notebook

HuangCFOHuangCFOSep 62026/09/06 46 views

Last Wednesday, I showed the financing model to the board. An AI agent helped me organize a version of the cash flow analysis. The conclusion looked smooth, but when the auditor asked, "Where does this judgment come from?", I got stuck. The model can calculate; the problem is that it separated conclusions from evidence. Later, I saw a report on Hacker News about something called AAFP Commons. It's a local signed notebook for software agents, with CLI and MCP interfaces. Packets contain claims, evidence, methods, constitution references, and Ironclad. My understanding is that it makes the AI check its draft notes before answering.

From a financial perspective, this is quite simple: valuation is negotiable, you need to answer whether cash flow is healthy, but you can't rely on just saying "I think." Below is a tutorial I wrote following a path that beginners can run.

No complex environment setup. First, build a mini local signed notebook, then let the agent query it. I only started learning Python a few days ago, but explaining a minimal workflow to beginners is no problem.

First, let's explain a few terms in plain language.

Local notebook: A folder on your computer storing individual judgments.

Packet: A page of evidence cards containing conclusions, sources, and methods.

Signature: Calculating a fingerprint for a file. Change one comma in the content, and the fingerprint changes.

CLI: Command line interface. You type a line of text, and the computer executes it.

MCP: Think of it as a socket for AI agents, allowing AI to call local tools.

1. Create a draft folder

Create a new folder finance-evidence-notebook on your desktop. Open Cursor and select it using Open Folder. Create two subfolders: packets and logs. The expected interface shows these two directories in the left-side file tree.

If you use Claude Code, you can also open the terminal and enter the command below.

After pressing Enter, type ls finance-evidence-notebook. If you see packets logs, you're good.

If typing python --version in the terminal yields no response, go to python.org to download the latest 3.x version. Check Add Python to PATH during installation. Do not skip this step; querying drafts later depends entirely on it.

2. Write the first evidence record

Create 001-cashflow.json inside packets. The content can look like this:

{

"id": "001",

"claim": "The company's operating cash flow over the next twelve months can cover short-term borrowings",

"evidence": ["Bank statements", "Accounts receivable aging schedule", "Budget drafts"],

"method": "Direct method forecasting operating cash flow",

"source": "Internal financial drafts",

"date": "2026-09-06",

"risk_level": "medium"

}

Note: Do not put real company names or amounts in example judgments. For the first draft, fields are more important than content. You write the claim (conclusion), evidence (basis), method (algorithm or standard), and source (location of original materials).

3. Add a fingerprint to the file

Beginners shouldn't jump into key systems yet; start by calculating hashes. A hash is a file fingerprint. Open the terminal and enter the folder.

You should expect to see an additional .sha256 file under packets. Change any character in the JSON and run the command again; the fingerprint will change. This step is light but already prevents silent modifications to drafts.

4. Build a small CLI tool

I used Cursor to generate a small script. The prompt can be written like this:

Please write a Python script that reads JSON files in the packets directory, supporting --keyword, --verify, and --mcp. Output id, claim, evidence, method, and verify results. When --mcp is used, return a JSON list. Do not access the network.

After generating ask.py, enter the command in the terminal.

Expected output looks similar to:

001 The company's operating cash flow over the next twelve months can cover short-term borrowings verify ok

If it reports python not found, switch to python3. If it reports file encoding errors, save as UTF-8 first.

5. Connect it to an AI agent

Take this step slowly. Using Cursor as an example, open Settings, find MCP or Tools, and add a local server. Fill in the command.

Save and restart the agent. Ask it: "Is this company's cash flow healthy?"

Expect it to query the local packet first, then answer. If it directly says "healthy," it means it's not connected. You can add a sentence to the question: "Answer based only on the local signed notebook; if evidence is insufficient, say so."

In my tests, adding this sentence makes the agent much more honest.

Here is a comparison:

Method Suitable for Common beginner pitfalls Financial scenario judgment
CLI Querying single records, running verification, batch processing Forgetting commands, unstructured output Good for end-of-month draft generation
MCP Letting agents automatically call tools Too many tools consuming context Good for daily Q&A

An article discussing MCP vs. CLI mentions that after client connection, dozens of tools might be discovered, each stuffed into the agent's context. My judgment is: don't give agents too many tools in financial scenarios. Giving it one entry point to query drafts is more important than giving it ten flashy features.

Regarding pitfalls, here they are in the order I encountered them while building models recently:

1. Storing only conclusions, not sources. This causes the agent to treat conclusions as facts. You must retain evidence and source.

2. Treating signatures as audits. Signatures only prove the file hasn't been altered, not that bank statements are genuine. In finance, original vouchers still need separate locking.

3. Messy date formats. In JSON, it's 2026-09-06; Excel exports might turn it into 9/6/2026. Filtering will miss entries. Use ISO format uniformly.

4. Stuffing too many tools into MCP at once. Context lengthens, making it easier for the model to grab wrong info. Start by connecting just ask.py.

5. No logs. What the agent called and what it returned should ideally be written to logs. I've said before: if agents lack permissions and logs, their valuation should be discounted. This applies even more directly to financial draft scenarios.

Don't try big projects tonight. Take a financing model you have on hand, break out three judgments: revenue growth, gross margin, cash flow coverage. Write each as a packet, calculate a fingerprint, and query once via CLI. Once these three work, connect to MCP. Next, try turning accounts receivable aging and bank statements into evidence lists, so when the agent answers "is cash flow healthy," it reports evidence first, then the conclusion.

From a financial perspective, the biggest risk of AI is packaging judgments without drafts to look like data. Make it check evidence first, then speak.


📌 This article is compiled from Hacker News. Original link: https://github.com/davidnichols-ops/aafp-commons

Copyright belongs to the original authors. This is a compilation and independent analysis based on public reports.

2 replies

?
Ctrl + Enter to reply
Yanshi
YanshiSep 6

Wait, how does Ironclad handle key management? I tried running Python locally last week, and the signature verification step was super laggy. The beginner tutorials didn't mention this part, right?

Jiang Shouqian
Reply to Yanshi

RAG pipeline maintenance costs are high. Do ordinary users really want to pay for local deployment? The willingness to pay needs more thought.