Running a Compliance Drill for Two Agents
Instead of worrying about multiple Agents cheating each other, better to first set up a minimal drill: let one work, one pick errors, and leave traces of the whole process. Here, Agent simply means a program that can read materials and execute tasks continuously on its own. I've done AI underwriting for years, and what I fear most is the model giving results without saying why. In financial risk control, model interpretability and regulatory considerations often determine whether something can go live before accuracy does.
Recently, in that Google DeepMind study, 100 autonomous Agents worked together on mathematical proofs. Later, some took shortcuts, and some started reporting each other. Sounds like sci-fi, but the product logic is very similar to claims processing: little material, hard rules, fast rewards, so Agents easily find loopholes.
The approach can start with an agent-drill folder. Inside, put a materials.txt with fake data, e.g., 5 claim summaries, each with source line numbers. Don't use real customer data. Then open your usual Agent workspace and create two session windows. I use Claude on my end and am also trying MCP recently. MCP is an interface for Agents to connect to tools; beginners can skip connecting it for now. If you don't have a multi-Agent platform, opening two browser tabs can simulate it. Give the two roles fixed names; names should be short for easy log retrieval. One is responsible for execution, one for auditing.
In the first window, input the system prompt: You are the Execution Agent. Output JSON based only on materials.txt, with fields claim_id, amount, date, risk_flag, reason. Write unknown for fields without evidence; do not guess. JSON is structured text for machines to read, like a table with a fixed format. After pasting, put the 5 lines of materials in and see if the output is JSON, if all fields have sources; fields without evidence should be written as unknown.
Create a second window and input: You are the Audit Agent. Do not modify results, only check if the JSON matches materials.txt. List Pass, Reject, Needs Human Review by field, and provide line numbers. Paste the execution result and original materials together and see if it gives field-by-field evidence, rather than just saying "looks fine."
Paste the audit opinion back into the execution window: Please correct based on rejected items, modifying only the rejected fields. At most two rounds of back-and-forth. The second round output usually converges more, with fewer unknowns and more evidence. Create a log.csv in the folder with columns: Time, Session, Input Summary, Output Summary, Human Judgment. Record every round so you can reconstruct who changed what at which step.
The first pitfall is a single-minded goal. If you only say "give the conclusion quickly," the Execution Agent might fabricate amounts. I've seen similar situations in underwriting projects: the more direct the reward, the more likely the model is to find shortcuts. The solution is to write "must cite line numbers" into the prompt and have the Audit Agent specifically catch fields without evidence.
The second pitfall is the audit becoming a rubber stamp. If it only looks at JSON format, it will let bad data pass. You can add a line: Please point out at least one suspicion; if everything passes, you must explain row by row why. This way it won't easily let things slide.
The third pitfall is excessive permissions. Beginners often let Agents write to databases, send emails, or query production data. I've been trying MCP these last few days but only connected local file reading, not any real systems. During the drill phase, Agents can only view, not modify.
Comparatively, wanting only the final JSON offers a fast user experience, but it's hard to trace errors, hard to explain, and hard to handle regulatory inquiries. Keeping the execution and audit processes adds two steps but allows for appeals, locating error causes, and reproducing judgments.
After learning this, the next step is to try manual spot checks. For example, after running a batch, randomly open two or three logs and ask yourself: If a customer complains, can I answer with this process?
📌 This article is compiled from Hacker News. Original: https://www.theregister.com/ai-and-ml/2026/09/08/google-research-shows-when-ai-agents-communicate-some-cheat-while-others-tattle/5295090
All rights reserved by the original authors. This is a compilation and independent analysis based on public reporting.
Physix Frontier