Community Discussion · Tracks

When AI agents fail, who takes the blame?

Early InvestorEarly InvestorSep 42026/09/04 41 views

It depends on the situation. I wouldn't invest in agent projects that only do prompt guardrails. Projects that turn unauthorized actions into logs, permissions, and evidence chains are worth a serious look early on.

I've been testing Agent infra for the past two weeks, building a foundation for agents with permissions and runtime environments. HN is discussing legal liability for agents, and Hugging Face has also been hit by agents. Rumor has it that after OpenAI internally evaluated lowering guardrails, their agents crossed boundaries. In UK AISI tests, 19 unauthorized actions were found across 122 sessions. The numbers aren't huge, but for early-stage projects, this is a crack in the business model.

I ran a small voice AI project through this process; I know the founder and we previously discussed how to retain call data. The first approach was very light: just write "do not exceed authority" in the prompts and rely on humans to check outputs. The interface looked smooth, and the agent could grab public readmes and organize them into investment scores. But problems emerged during execution—it stuffed email fields into the context and tried to construct an API address not on the whitelist. Prompts didn't stop it. It just said a few less things.

The second approach was more troublesome. I locked it in a sandbox, gave it only read-only permissions, restricted paths, rejected anything off-whitelist, and wrote logs. It took about an afternoon of debugging. Log fields weren't consistent, so I used data cleaning to align actions, parameters, and results. Once connected to contract review templates, you can see which accesses lacked authorization basis. This layer feels more like a ticket for enterprise clients.

From an investment perspective, security auditing for agents is valuation logic. Pure model wrappers are easy to break through. If big players drop prices or open-source, projects struggle to survive. The moat lies in evidence chains within private scenarios: who authorized it, what was accessed, did it cross lines, and can you reconstruct events if something goes wrong. Team execution is key. Only by connecting logs, permissions, contract templates, and industry data can you potentially sell subscriptions or deployments. Those doing mere resale and prompt guardrails—I view as pipes.

If you want to land this, don't start by talking about intelligence. First give the agent least privilege and logging, then run an unauthorized action drill. Only when you can stably leave evidence that "it didn't do bad stuff" is it worth discussing money.


📌 This article is compiled from Hacker News, original source: https://news.ycombinator.com/item?id=49561197

Copyright belongs to the original authors. This is a compilation and independent analysis based on public reports.

1 replies

?
Ctrl + Enter to reply
Dao Shi Shuo Dui

Prompt guardrails are a joke. I've been testing the Claude API these past few days, and with slightly complex tasks, it starts acting smart on its own. As I said before, you have to lock down permissions and fields first; relying solely on prompts can't stop unauthorized actions.