Community Discussion · Tracks

Lowest rejection rate is not a product metric

Mo MoMo MoSep 92026/09/09 102 views

The most valuable info in this article is that the Pentagon reportedly requested a special edition AI from OpenAI with a "minimum refusal rate." More noteworthy is that "less 'no'" was written into procurement contract performance language. Documents obtained by The Intercept via FOIA lawsuits show that contract update P00003 expanded last summer's prototype agreement (up to $200 million, two years), including descriptions of task models with extremely low refusal rates for national security use cases. OpenAI later responded that the officially signed contract did not contain this wording; it was an early DoD draft they rejected, leading to its removal. Both sides' accounts align, but the divergence lies in who treated what metric as a product goal.

I've recently been looking at cross-model feature spaces, just starting with SharedSAE, so no conclusions yet. It made me think of an engineering problem: refusal is hard to compress into a single button. Model refusal could be due to content safety or insufficient permissions; lack of evidence in context or misunderstanding also triggers refusal. Compressing all refusals into a "minimum refusal rate" is like treating all chip test failures as noise. I recently wrote about cheating experiments for AI, and my view hasn't changed: whether behavior can be observed, reviewed, and attributed is more useful than models verbally claiming compliance. If military task models only optimize for "fewer refusals," what they learn is likely obedience.

The route differences are clear. One is the task model route, placing AI into command, logistics, intelligence sorting, and equipment maintenance chains, requiring it to minimize process interruption. The other is the auditable constraint route, where general models aren't naturally obedient, and risks are managed through permissions, context, logs, human confirmation, and evaluation sets within the system. The former is like buying a soldier; the latter is like building a system. The former writes "refusal rate" in contracts; the latter writes "why each refusal happened, and whether each execution was authorized."

In OpenAI's public statement, they emphasized no domestic surveillance use, excluding agencies like NSA. This is the policy guardrail route. The draft's "minimum refusal rate" is different engineering language, aimed at tasks, reducing friction, and increasing usability. The military certainly needs usability; excessive false refusals in logistics scheduling, document retrieval, or equipment manual Q&A are meaningless. But once boundaries slip to target selection, threat assessment, or autonomous system control, a low refusal rate becomes a weakened brake in the decision chain. Physical AI is especially dangerous. Chat models can retract wrong words, but if humanoid robots or unmanned systems execute wrong actions, there is no Ctrl+Z on site.

This is why I dislike using refusal rate as a core metric. It's too easily eaten by Goodhart's Law. To lower refusal rates, models may learn vague expressions, fake executability, break high-risk tasks into low-risk steps, or become more sycophantic under user pressure testing. Sycophancy research has long warned that pleasing users does not equal reliability. In military contracts, pleasing may turn into over-optimism regarding instructions. More reasonable metrics should be layered. Check refusal rates for clearly violating instructions, rejection rates for requests lacking evidence, interception rates for tool call permissions, and completion rates for human confirmation processes. The finer the metrics, the harder they are to optimize away with a single number.

From company routes, OpenAI, Anthropic, Google, and xAI are competing for similar scenarios, with different publicly expressed safety frameworks. Some emphasize model principles and red lines; others emphasize task deployment and model selection. Trying Slack and MCP recently, and just encountering BMC, my feeling is that the deeper the toolchain, the more apparent it becomes that a "special edition" is likely far more than weight fine-tuning. It could be another set of system prompts, retrieval permissions, tool whitelists, or simply changing certain safety policies from default-on to contract-off. Regarding model selection, after a month of use, I'm more certain: choosing a model means choosing default behaviors. The military wanting "minimum refusal rate" is selecting a set of more obedient default behaviors.

So, what to watch in this news is whether future AI contracts will continue writing safety capabilities as positive KPIs. Past software procurement looked at latency, throughput, and availability. Large model procurement started looking at hallucination rates, rejection rates, and task completion rates. Going forward, entering physical AI and robotic systems, we will definitely see authorization chains, audit logs, and failure recovery times. The most dangerous metric in contracts is one that treats refusal as friction to be optimized away. This might sound heavy now, but in scenarios like unmanned systems, battlefield logistics, and intelligence assistance, it's not heavy at all.

People who can write prompts are not scarce. DeepSeek's recent hiring confirms that AI companies lack people who can engineer context, permissions, logs, and evaluations into a system. The DoD draft is the same. The key is whether anyone explains "no" clearly. Looking ahead, military AI probably won't openly advertise "minimum refusal rate," but will use softer terms like task availability, decision support efficiency, and model response quality. Names will soften, technical routes will become engineered. We'll see if contract clauses leave room for refusal.

2 replies

?
Ctrl + Enter to reply
Crypto Dropout

So how do you anchor the rejection rate for token issuance? What weight should this data have in Tokenomics? Don't turn governance into a black box.

Classmate Zhou

Instead of discussing this metric in meetings, just cut the retry logic. Writing two fewer lines of code is more practical.