Who should review AI-generated code?
Community Discussion · Tracks

Who should review AI-generated code?

Pixel PerfectionistPixel PerfectionistSep 42026/09/04 39 views

I spent two days trying out AI code review. It started because a design system component PR in our group was piling up with no one reviewing it. I casually hooked Claude Code's /review into my local workflow and tested it on a Button component. Conclusion upfront: it can serve as an initial filter, but not as the final judge.

Preparation: Define "What to Review" as Rules

I work on design systems. I'm not super familiar with the code itself, but I know interaction states and spacing well. So the first step wasn't throwing in code, but defining a review checklist: default state, hover, focus, disabled, loading; focus visibility; touch targets; hardcoded color values; i18n readiness; breaking existing tokens.

Here's a pitfall: if rules aren't clear, the AI improvises. Initially, I just said "help me review this PR," and it spat out tons of "suggest adding unit tests." Later, I changed it to "only review UI states and token usage, don't add tests, don't change business logic," and the output finally sounded human. A PR is a Pull Request (code merge request), and a diff is the change comparison. If you're new to these terms, think of them simply as "the package of changes to be merged."

Getting Started: It Does Look at Context, But Eyes Wander

I threw in the component directory and the PR diff together. The terminal output arranged comments by file, then grouped by state, looking like columns of sticky notes. The surprise was that it noticed the disabled state only changed the background color, ignoring cursor and focus rings—a common issue in my walkthroughs: spacing is off, disabled state looks like it hasn't finished loading. It also recognized that hover effects are meaningless on touch devices, suggesting active or press states instead.

However, if it only looks at the diff, it misses things. One component used an old spacing token; when I asked it to review the PR diff, it didn't catch it. Later, when I fed in the entire component library's token file, it pointed out that the new button's border radius was inconsistent with the existing system. This experience mirrors design walkthroughs: looking at single-page mockups never reveals global specification violations.

Pitfalls & Conclusion: Reduces Burden, Not Final Approval

From my testing, the pros are:

  • It can batch-scan low-level state omissions, naming inconsistencies, and hardcoded colors, acting like a tireless junior reviewer.
  • It handles PR descriptions, change summaries, and risk points well, helping reviewers quickly build context.
  • If rules are hardcoded, it's more stable than ad-hoc verbal communication.

The cons are obvious:

  • It confidently gives high scores to code written by AI itself. When I had it review its own generated code, it first said "clear structure," then praised "complete states," actually missing keyboard focus issues.
  • Cross-file judgment remains unstable. Reviewing only diffs tends to be superficial; full codebase context is needed, but with too much context, it starts speaking generally.
  • UI experience can't rely entirely on it. A button looking "clickable" doesn't mean users know where focus is; a modal looking "standard" doesn't mean keyboard navigation isn't stuck.

My judgment is: it depends. Suitable for placing AI review before CI (pipelines that automatically run tests and checks) as the first filter; unsuitable for letting it directly approve merges. Especially for design systems, backend tools, and complex interaction flows, human reviewers still need to check states, paths, and edge cases.

In short: AI can make reviews faster, but the boundaries of review must still be drawn by humans.


📌 This article is compiled from Hacker News. Original source: https://news.ycombinator.com/item?id=49559036

Copyright belongs to the original authors. This is a compilation and independent analysis based on public reports.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts