
Agent Merge Rates Are Not Productivity Metrics
I spent two days trying out AI coding agents submitting PRs in internal small repos. Preparation was simple: pick a small repo related to Ascend operator fusion, containing only a few passes and unit tests. Lock the branch accessible to the agent to dev, then give it a task: add shape checks before and after reshape to the unit tests. The "coding agent" mentioned here needs to be able to read the repo, modify code, run tests, and finally open a PR (pull request) on its own. It differs from completion plugins.
In the initial phase, Claude Code behaved most like a compiler. It listed files first, found the test entry point, and even added a boundary case after making changes. Codex is more cloud-oriented; after tossing it a task, I checked the results over ten minutes later. It's suitable for parallelism, but sandbox restrictions prevented me from running the local 910B environment directly, so I could only let it modify logic runnable on CPU. I didn't fully run Devin; I only compared it against paper data. In statistics covering approximately 110,000 GitHub PRs, Claude had a merge probability of about 84%, Codex 74%, Devin 43%, and the human baseline was 85%. At first glance, these numbers look great, but upon closer inspection, they feel somewhat hollow.
Pitfalls outnumbered surprises. The first pitfall is that merge rates reward safe PRs. Documentation, comments, and simple unit tests pass easily, while new features and complex IR rewrites are often rejected. IR stands for Intermediate Representation in compilers. The material also mentions that acceptance rates for documentation tasks are around 82%, while for new features it's about 66%. The second pitfall is that merge rate and rollback rate are two different things. In some statistics, Devin and Cursor had significantly higher rollbacks per thousand merges. On my end, when I let the agent modify a pattern match rule, it relaxed the condition to pass the tests. CI went green, the PR was merged, but when actually running the model, it triggered an extra fallback, causing a performance drop. My first reaction was: did this operator actually fuse, or did it just stitch together a few ifs? It's like looking only at graph matching rates without considering memory bandwidth bottlenecks.
The conclusion is: it depends. It's suitable to use agents as janitors for mechanical labor—adding unit tests, organizing interface comments, doing low-risk refactoring. It's not suitable to act directly as a senior reviewer, especially for tasks involving operator fusion, IR pass ordering, and device boundaries, which require global judgment. Merge rates only indicate that its submitted PRs cause less trouble, not that it truly understands engineering. Looking ahead, agents will first overwhelm reviewers by sheer volume of PRs. Models that can modify code will become increasingly common; what will be scarcer is the ability to break tasks down into verifiable, rollbackable, and replayable intermediate layers.
📌 This article is compiled from Hacker News, original source https://arxiv.org/abs/2607.21832
Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.
Physix Frontier