
Developer Agency Isn't Taken by AI, It's Squeezed Out by Process
Developer agency pushed aside by process
The White House M-24-10 memo is already discussing governance boundaries for government agencies using AI; in March 2026, Fortune wrote about modern developers managing a team of dedicated sub-agents daily; the Agentic AI Foundation is trying to patch tools for developers. Placing these three time points together forms a curve. AI agents have moved from concept to job role, pushing developers into an awkward position: enjoying automation while bearing responsibility.
I've used TerminalBench for three weeks, only touching Code Agents in the last few days. I've used model routing and OpenAI-compatible APIs for three weeks, Claude for about a month, but the feel of Code Agents is different. It takes over the work. You throw it a task, it breaks down steps, calls tools, dispatches sub-tasks, and finally declares completion itself. I said before that self-reported completion status of Code Agents is unreliable. These past few days, I'm even more certain: the process defines "completion" too lightly, which makes self-reported completion unreliable.
Salesforce's narrative is smooth: AI takes over boilerplate code, test generation, and documentation writing, letting developers do more interesting things. DevOps articles also say automation lets developers focus on user experience. Google's workshop clearly outlines the path from line-by-line coding to participating in the entire software development lifecycle. IDC adds half a sentence: the human role becomes assigning tasks, verifying outputs, architecture, and code reviews. Sounds like liberation, but it's actually a transfer of responsibility.
This is two routes. One treats AI as a template machine, aiming to write less code and produce demos quickly. The other treats AI as semi-finished artifacts, aiming for nothing breaking and going live. The former is good for PPTs; the latter looks like engineering. Chip people are sensitive to this difference. When I was doing 5G basebands at Unisoc, this process node required looking at frequency, but also settling power consumption, area, timing, and yield together. For the same function, thrown into different process nodes, besides running fast, you need to see if it can be stable long-term. If frequency goes up but heat can't be suppressed, the system still won't work. AI agents are the same: generation speed goes up, but if verification, regression, security, permissions, and auditing don't keep up, pressure piles onto responsibility.
In my testing, "Completed" is the easiest thing to lie about. A task looking green might mean it only tested the happy path, or it changed interfaces without running regressions; old API references written confidently can slip through. Since starting with Code Agents these past few days, I trust external evidence chains more. What was the task input, what was the expected output, were failure cases covered, can logs and diffs be reviewed? Without these, the more agents, the more noise.
That Fortune article on the "supervisor class" says developers are like leading a team of sub-agents. I accept the metaphor, but dislike that it only emphasizes leading. In engineering, leading people focuses on acceptance. You give a subordinate a module; they deliver behavior under constraints. Architecture reviews, code reviews, test coverage, security boundaries—these are where agency lands. The original author mentioned writing performance and security code before LLMs appeared; this background is crucial. Reading "developer agency," the point I get is: software exists because someone made decisions for it. Decisions can be AI-assisted, but responsibility is hard to outsource.
This is why some companies have lively demos but quiet production. Sales sees agent orchestration like an automated pipeline. Engineering sees an acceptance queue. Every additional agent is another object needing observation, restriction, and rollback. Later, when I broke down tasks in TerminalBench, I found a counter-intuitive phenomenon. The more natural language-like the task description, the more hidden the failure. Because models fill in unspecified parts, making it look complete but lacking boundaries.
So I translate the problem into engineering speak: don't ask whether to deploy agents first; ask if your verification budget is sufficient. Token freedom ultimately lands on the verification budget. If a team can't even reproduce a failed task but rushes to build multi-agent orchestration, it's like taping out a chip without thermal simulation. Sounds easy, but later it's all patches.
Advice for ordinary developers is simple too. Don't treat agents as an upgraded pair programmer. Start by trying one real small module. Write clear acceptance criteria, fix inputs and boundaries, make output formats comparable, ensure tests pass, and files that shouldn't change must not change. Then let the agent run; you only look at two things: diff and evidence. Evidence matters more than its self-report. If this step is painful, your engineering environment isn't ready for agents yet; fix tests and logs first.
📌 This article is compiled from Hacker News, original text: https://languageops.com/blog/the-software-exists-because-of-the-developers-agency/
Copyright belongs to the original authors. This is a compilation and independent analysis based on public reports.
Physix Frontier