
Check for documentation before letting AI contributors into the repo
Recently, I've been looking into specification issues before AI coding agents enter repositories, and I tried RepoPolicyScore. I threw a small lab repository containing annotation scripts into it to see if it would tell me what's missing when AI coding agents come in to modify code. So-called AI contributors are tools like Codex, Claude Code, and Cursor; they read code, modify code, and submit PRs based on instructions in the repo. The interface entry is straightforward: an input box prompting you to paste a public GitHub repo URL. After pasting it, I waited over ten seconds, and the page started running checks. Officially, there are 25 checks examining files like contribution docs, README, issue templates, and PR templates. Each item provides file and line numbers, resembling reviewer comments—it directly points out, for example, that CONTRIBUTING.md line X doesn't specify how to run tests.
Your repo never told them what AI-assisted work is allowed.
This hit home. Previously, I wrote about hiring AI; once resumes are templated, systems tend to optimize for single metrics. Now, code contributions face similar risks. AI agents reading repos read files. If things aren't written clearly, they might guess based on their own priors. I've used Claude Code and Codex for about a month and found that agents fear broken context the most; code complexity is secondary. This tool acts like a health check for repos, verifying if executable instructions have been left for machines.
The advantages are obvious. It breaks down vague concepts like "AI-ready" into checkable items, so maintainers don't have to audit PRs by gut feeling. Reports locate specific files, saving the hassle of flipping through everything. Methodologically, it assesses auditability, asking if machines can know where the boundaries are once they enter. For labs with tight budgets and student code often lacking documentation, it serves as a reminder. While this direction alone might not yield publishable papers, the auditability of contribution policies is a solid topic. There are also scorecards in the community based on frameworks like OpenAI Agentic Legibility; RepoPolicyScore seems to be the branch leaning towards policy compliance within the same lineage.
The disadvantages are significant too. It feels more suited for checking public repos. When I first pasted an internal fork address, no results came up; I switched to a public mirror to get it working. If a project originally lacks a CONTRIBUTING.md, the report might count many blanks as issues, making it look alarming. It also struggles to judge business risks; for instance, if annotation processes involve privacy but it's not written in the files, it might not realize the severity. More troublesome is that such checks can induce teams to fill in template jargon rather than genuinely clarifying contribution boundaries.
In my testing, the biggest surprise was its hint that my PR template didn't require stating "whether modifications were AI-assisted," while the README said demos could run automatically. This indicates contradictions between documents; humans can infer meaning, but machines might not.
The real question is whether the repo clearly defines contribution boundaries. So, I recommend open-source maintainers, mentors, and anyone having students submit code try it out, at least to know which rules are unwritten. I don't recommend it for individual users who just want AI to quickly generate PRs; it primarily helps you point out pitfalls in advance.
📌 This article is compiled from Hacker News, original text: https://repopolicyscore.com
Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.
Physix Frontier