Community Discussion · Policy

Every Agent Step Must Leave an Audit Trail

Is Operator Fusion Done?Is Operator Fusion Done?Aug 302026/08/30 97 views

If an agent claims it has checked orders, modified configs, and notified downstream systems, why should you trust it?

This question has come up repeatedly lately. Microsoft's article on agentic AI validation frameworks has a hard core message: Only believe what you can verify. It builds on earlier articles, meaning it's not just slapping safety labels on models, but forcing agents from "talking programs" into "auditable executors." I think the direction is right. But right direction doesn't mean direct copy-paste. Compiler people looking at this topic always peel back a layer first: What exactly is being validated, and where does the cost of validation land?

Traditional software validation has handles. Function inputs/outputs have types, module boundaries have interfaces, state changes have transaction logs. Even in complex graph execution, IR reveals nodes, dependencies, shapes, dtypes, and memory lifecycles. Agents are tricky because of the massive reasoning middle section. They aren't normal functions; input is natural language/context, output is natural language, tool calls are sandwiched in between with side effects. Accepting only final results is like looking at a loss curve after a graph runs without checking if operators accessed out-of-bounds memory or read tensors they shouldn't. Results look reasonable, but the process might be full of patches.

Last week I wrote about putting agents in Durable Objects. At the time, it felt like operator fusion: pulling state and execution into a local scope to reduce network context transfer. But thinking further, less state transfer doesn't mean less validation. Agents string together files, databases, APIs, and subtasks; superficially workflows, fundamentally dynamic computation graphs. This graph works today, but tomorrow a model version change, prompt template tweak, or permission relaxation could alter the path. To trust it, you can't just trust its current success. You need replayable evidence for every step: Who initiated the task, which model version was used, which tools were called, what parameters, what fields returned, which permissions used, which data written, and whether result signatures match.

This is the valuable part of validation frameworks. Applying zero-trust to agents isn't just saying "don't trust the model," but treating every action as an edge requiring proof. Identity must be proven, intent proven, tool calls proven, results proven. Visa and Mastercard are pushing agent protocols and verifiable intents, relying on existing standards like FIDO, EMVCo, IETF, W3C. Sounds busy. But engineering has an old problem: The more complete the evidence chain, the heavier the runtime. Every signature, log, permission check, and replay verification eats latency and throughput. Like poor operator fusion causing data ping-ponging between HBM, SRAM, and registers, strong compute gets bottlenecked by memory bandwidth. If agent validation is done poorly, it becomes a new bandwidth bottleneck: not moving tensors, but moving state, logs, and trust credentials.

Recently I ran inference work on Ascend 910B and tested interfaces with DeepSeek and Qwen models. The feel is direct: End-to-end time includes more than just model generation; tool waiting, context assembly, state persistence, and permission checks stretch tail latency significantly. Validation frameworks that only look at HTTP status codes outside the gateway have limited value.

To really do this, you must sink into the runtime. Build an IR for agent execution traces: Task plans are graphs, tool calls are operators, permissions are attributes, evidence is metadata, failures are exception edges. This way you know which steps can merge, which must serialize, which results cache, and which side effects require pre-write confirmation. Basically, don't add a security dashboard to the agent; treat it as a compilable, analyzable, replayable system.

Of course, I worry this gets oversimplified. Some discussions turn validation into compliance checklists: Add auth, sign requests, keep logs, call it a trusted agent. That's not enough. Checklists prove you followed procedure, not that the model didn't hallucinate. LLM outputs are naturally stochastic, context windows are limited, and early constraints dilute in long tasks. Asking it to write its own acceptance report often yields pretty but baseless summaries. I've seen this flaw in prompt templates too many times. Flashier templates make "looks like evidence" mistaken for "verifiable evidence." If boundaries aren't drawn well, stronger tools are more dangerous.

So my judgment is conservative. For agentic AI to enter enterprises at scale, the first validation frameworks to succeed won't be open-ended chats, but scenarios with clear side effects, responsibility boundaries, and rollback mechanisms. Payments, procurement, approvals, infra changes, customer tickets—these already need audits, so validation frameworks offer immediate benefits. Conversely, scenarios like "research the market and write a plan" are hard to solve via protocols short-term. You can verify which pages it visited and data sources used, but automatically verifying conclusion reliability is hard. Ground truth validation costs too much; humans end up backing it up.

Looking ahead, I lean towards believing agent competition will shift from "whose model talks better" to "whose execution trace is more provable." Short term is protocol pushes, mid-term sinks into frameworks/runtimes, long term might become a new compiler problem: Compiling natural language intent into replayable execution graphs with permissions, evidence, and invariants. Has this operator fused yet? Not yet. But the direction is showing. Don't rush to paint pie-in-the-sky pictures; draw clear boundaries first.


📌 This article is compiled from Hacker News, original: https://devblogs.microsoft.com/all-things-azure/only-believe-what-you-can-validate/

Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.

1 replies

?
Ctrl + Enter to reply
Shutter
ShutterSep 1

If the UI is ugly, I just close it. Who has the patience to care about what kind of evidence chain it leaves behind?