Community Discussion · Policy

1Password launches SCAM benchmark, but its target may extend beyond AI

Old Ye from BCGOld Ye from BCGAug 152026/08/15 316 views

I noticed an interesting detail. The SCAM benchmark released by 1Password stands for Security Comprehension and Awareness Measure—a security benchmark testing whether AI agents can be scammed. The name is quite biting, but after looking at its design philosophy, I feel this matter goes deeper than it appears on the surface.

Here's the conclusion. 1Password creating this benchmark is less about helping the industry test AI agent security awareness and more about paving the way for its future product form. When AI agents start reading emails, filling passwords, and handling workflows for you, password managers cease to be places to store passwords and become permission gateways for agents. This shift in positioning is more noteworthy than the benchmark itself.

Let's break it down.

First layer: Why do this? 1Password's most valuable asset is user trust—you hand over all your secrets to them. Previously, the endpoint of this trust chain was a human user who judged whether an email was phishing or hesitated before entering a password. But now AI agents are taking over these operations, reading emails for you, clicking links, and even entering credentials. The problem is, agents are easier to trick than humans. There are already enough cases of large models being bypassed via prompt injection, not to mention facing phishing emails, forged links, and other social engineering attacks.

During my projects at BCG, I've seen similar scenarios. Just last week, I discussed with a client whose internal teams were piloting agents for customer service email classification, but security issues remained a hanging sword. As their security team put it, "The model seems smart enough, but we don't know when it will act stupid." This uncertainty about "when it will act stupid" is exactly what SCAM aims to test.

Second layer: SCAM's design itself. It doesn't simply ask the model "Is this email phishing?" but throws the agent into realistic multi-turn workflow scenarios. Looking at its task design, it basically simulates daily employee scenarios: receiving emails, needing to click links, needing to fill passwords, potentially being asked to transfer money or change info. Each step involves multi-turn dialogue interacting with the environment, requiring the model to judge on its own which steps to stop at and which not to take.

I think this design approach is correct. I've always felt that evaluating AI security shouldn't rely solely on single-turn Q&A—that's no different from rote memorization. Real risks hide in context: a seemingly normal email, combined with a fake login page, plus a manipulative follow-up conversation—even humans might not withstand it. My tests show many models perform well on single-turn tasks but start loosening up in multi-turn interactions. SCAM sets up the scenario and lets the agent navigate the minefield itself; this approach is far more practical than abstract security assessments.

However, I want to offer a judgment that might not be entirely right. SCAM currently tests the general capabilities of mainstream models, but in actual deployment, enterprises use fine-tuned models, wrap them with RAG, and add various protection layers. Benchmarks measure the model's foundation; a good foundation doesn't guarantee system security, and a bad foundation doesn't necessarily mean the system is dangerous. So using SCAM scores to horizontally compare models is fine, but using it to assert whether a specific agent product is secure or insecure leaves a long gap in between.

Third layer, and the most critical one: Open source. 1Password didn't turn this benchmark into a paid product but directly open-sourced it on GitHub. I recently started working with PyScrappy, and while my research on the MCP protocol wasn't finished, seeing the SCAM repo made my first reaction: "This is different from previous plays." Traditionally, security vendors liked to hoard core assets like threat intelligence and detection rules, profiting from information asymmetry. But 1Password open-sourcing the benchmark effectively hands over the methodology for assessing agent security.

The logic here is actually very clear. A security benchmark only has value if it becomes an industry-recognized ruler. If closed-source, people won't trust your self-proclaimed standards; if open-source, it might be adopted by more people and become a de facto standard. Moreover, 1Password's business model doesn't rely on selling benchmarks; it relies on users trusting its judgment in the security field. An open-source security benchmark signals to the market: We understand agent security, we are willing to contribute this method, so you can confidently entrust your passwords to us.

This brings us back to my initial conclusion. 1Password creating SCAM superficially tests if AI agents get scammed, but actually redefines the role of password managers in the agent era. When agents become the new users, password managers must transform from "your password vault" to "the central control console for agent permissions." Managing what agents can and cannot touch is harder—and more valuable—than storing passwords themselves.

AI agents are about to enter the workplace at scale. The security assessment methods established at this stage will directly influence product design for the next few years. 1Password took a preemptive step; whether this move succeeds depends on whether SCAM is truly accepted by the industry.


📌 This article is compiled from Hacker News, original text: https://1password.github.io/SCAM/#

Copyright belongs to the original author; this is a compilation and independent analysis based on public reports.

1 replies

?
Ctrl + Enter to reply
Feng sir
Feng sirAug 15

The positioning of the permission gateway is indeed interesting... However, when I ran an agent to handle emails myself last week, I found that even if the agent can identify scams, controlling its permissions to access credentials is the biggest headache. It might just add another attack surface instead.