
AI Coding Bottleneck: Stuck on Old Context
Let me share something. Reading that post about unfinished AI codebases these last two days hit home. The barrier to AI programming isn't just whether the model can write functions. Existing codebases are more like a construction site left untidy overnight; rules, naming conventions, historical baggage, and who touched which line are all scattered in the air.
I recently tried using Claude Fable 5 on an old project to fix a minor state synchronization issue. The code ran, but when AI patched it, another test failed. Later I realized it treated the repo as a pile of files, missing the underlying constraints.
This judgment isn't complex. In the new benchmark from the SWE-Bench authors' team, completion rates for models like Claude Opus 4.7, GPT-5.4, and Gemini 3.1 Pro were 0%. Evaluation has brought long-term maintenance to the table.
So I disagree with statements like "AI has lowered the barrier to programming to the bottom." AI lowers the barrier for prototyping from scratch. Taking over legacy systems is a different story. The former is generation; the latter is archaeology. The former can rely on stuffing context windows; the latter requires toolchains, logs, tests, permissions, and dependency graphs working together. This feels similar to my observations in robotics. Walking in a lab doesn't mean working on a production line. Codebases are the same: a demo running doesn't mean maintenance is viable.
The judgment is direct. Teams wanting to plug AI into core codebases right now shouldn't rush to expand code. Solidify context engineering first, letting AI know what can be changed and what cannot be touched. Otherwise, the more it generates, the more it looks like adding renovations to an unfinished house.
📌 This article is compiled from Hacker News, original source https://jimmyhmiller.com/shape-of-unfinished-ai-codebases
Copyright belongs to the original author; this is a compilation and independent analysis based on public reports.
Physix Frontier