The hard part of AI software factories: who handles merges?
After reading this report on the AI software factory, my judgment is clear: getting an agent to open a PR isn't hard; what's hard is entering the real engineering pipeline. Behind the software factory lies repository structure, CI, permissions, rollbacks, security scans, and owner sign-offs. As long as these links aren't connected, automation can only stay at the demo stage.
The report title mentions that agents can open, review, and merge PRs. These three actions look simple but are actually three capability lines. Opening a PR requires understanding the issue, locating the module, generating a patch, writing the change reason, and ideally including tests. Doing a review requires reading diffs, judging interface breaks, boundary conditions, and security risks. Auto-merging into the main branch turns risk from suggestion into fact.
Templated solutions like Microsoft Foundry also illustrate this direction. Officially, voice agents, release management, and data unification are made into preset templates, bundling Azure and GitHub quick starts. But templates solve startup costs, not production responsibility.
Compared to Copilot, I've used it for about 2 months. It definitely saves effort on completion, boilerplate code, and small single-file fixes. But once you expand from issues to PRs, it easily generates things that look complete but haven't actually run. I recently started trying GPT and Claude, using them for less than a week. I tried the WuJie plugin for just 6 days; the PR descriptions look decent, but a pretty description doesn't mean the code is mergeable.
What gets stuck is the chain of responsibility and verifiability. The easiest place for AI coding tools to deceive people is when the output is too smooth.
I wrote a post earlier about AI slop; the core meaning is this: batch-generated content looks like human speech but is informationally empty. PRs are the same. An agent can write that it fixed a null pointer, added unit tests, and improved performance, but if no one can verify it, these sentences are just wrapping paper.
What's hard in engineering is verifiability. Did the tests run? Did coverage drop? Can the migration roll back? Was the permission model broken? Can logs pinpoint issues? None of these can be taken responsibility for by a casual sentence from the model. When we run small repos, the biggest fear is it changing unrelated files together. Long context forgetting makes completions look smart, but the code drifts. Tools like WuJie have merit in their completion logic, but long-context errors are more dangerous in PR automation, as one change might span multiple files.
Experience with WorkBuddy also reminds me that cleaning non-standard materials is unstable. Dirty data goes in, and results easily get messy. Code repositories also have dirty materials, such as outdated READMEs, invalid comments, historical issues, and semi-abandoned modules. If the agent lacks context weighting and dynamic judgment, it easily treats old specs as current facts. In enterprise selection, I previously cared more about model capability; now my thinking has changed. Whether it can be swapped, whether there is auditing, and whether the supply chain is secure are often more practical than the model being slightly smarter.
The Microsoft Work Trend Index Annual Report says 82% of leaders express confidence in using digital labor to expand employee capacity in the next 12-18 months. This ratio is not low. But this is leader confidence; the engineering chain of responsibility hasn't been connected yet. Whether the factory can run depends on who manages accidents.
The landing route can proceed from low to high risk. Start with read-only reviews. Let the agent comment in PRs, give suggestions, touch no code, and touch no merges. This stage is easiest for evaluating real judgment. Finding empty catches, pointing out migrations without rollbacks, and reminding about interface compatibility is more valuable than generating a hundred lines of code.
Next, do draft PRs. The agent opens a branch, generates code, runs CI, but the status must be Draft. Merging is only allowed after human confirmation. The key is shifting review from reading code beforehand to verifying evidence afterward, including test results, affected file lists, change explanations, and rollback plans.
Only lastly discuss auto-merge. The scope must be very narrow, such as documentation, config boilerplate, low-risk internal tools, pure UI copy, and formatting changes with no business logic. For core business processes, permissions, payments, data migrations, and security boundaries, don't expect full automation in the short term.
Predicting a trend. In the next 12-18 months, the AI software factory will likely popularize review agents first, then move toward semi-automatic PRs. Scenarios daring to auto-merge will likely be low-risk internal projects and highly standardized code; large production repos will lag behind. Once the factory lands, who signs off and who is responsible will be brought to the forefront.
📌 This article is compiled from Hacker News, original: https://www.firecrawl.dev/blog/ai-software-factory
Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.
Physix Frontier