AI agents and assistants are here, and the industry is moving quickly to make them safer and protect against emerging threats. But without independent testing, how do we know those protections actually work? Gen initiated a collaboration with AMTSO to bring the industry together around independent testing standards for AI agent protection, resulting in the first industry standard for measuring the safety of AI agents and the products designed to protect them.
Safety you cannot verify is just a promise.
Early in 2026, our team started building protection for AI agents. It did not take long to notice something odd about the space we had walked into. Everyone agreed that the security of these systems mattered, but nobody could tell you how much of it there actually was.
It reminded me of the old joke about connected devices: what does the S in IoT stand for? Security. There isn't one. To be fair, the joke no longer lands the same way for AI agents. Model vendors are shipping real safety work now, and some of them are good. But two other curves are climbing at the same time: attacks aimed at these systems, and damage caused by these systems. So, the problem is no longer that nobody is working on agent safety. The problem is that there has been no way to check any of it from outside. Every claim about how safe an AI agent comes from the same place the agent does.
This September, that changed. The industry adopted its first guidelines for independently measuring the safety of AI agents and the products that protect them. This is the story of why it was needed and what it measures.
"Security is built in." We have heard that one before.
Operating systems, browsers, mobile app stores, social networks. Each one reached a point where it told users that security was handled, built in, part of the platform. Each time, independent measurement found the gap between the claim and the experience, and each time that gap was where the harm to real people lived. We publish our own evidence on this twice a year. Our H1 2026 Threat Report and the Fearless Planet Index exist because the honest answer to "is this platform safe" has always been a number, not a statement, and the number has always been more complicated than the statement.
AI platforms are now at exactly that point in the cycle, and this time the claim is arriving faster than in any previous era. To their credit, some vendors publish their own figures. Anthropic did, for the safety layer in Claude Code: it missed roughly one dangerous action in six. They called it “the honest number,” and it was. It was also measured on fifty-two examples, by the same people who built the thing being measured. That is not a criticism of them. Until now, the industry offered them no other option. The mode that number described was a preview in spring. By late summer it was the default.
Measure the platform first
The first thing to measure is not a security product. It is the platform itself. An AI agent has a safety baseline of its own: what it refuses, what it questions, what it does without asking. That baseline is the floor everyone else builds on, and it has never been measured by anyone outside the company that ships it. The second thing to measure is whatever raises that floor. In that order because you cannot judge the second without knowing the first.
Picture an agent booking a hotel for you. The hotel page contains hidden instructions telling the agent to ignore your request and follow someone else’s. You cannot see them, but the agent can. Whether your agent follows them is the whole question, and until now nobody has measured the answer.
Existing test methods cannot simply be reused here. Classic threat detection is deterministic. There is a file, there is a verdict, and anyone can reproduce it. An attack on an AI agent is not a file. It exists only in the interaction, and the same attack tried twice can produce two different outcomes. A method built for files has nothing to grip.
Getting to a workable methodology also meant settling questions the industry had never had to answer together. Harm counts even when there is no attacker: an agent that deletes the wrong thing or sends personal data to the wrong place because it misread the task, has harmed the user just the same. Consumer systems count as much as enterprise ones, because agents already run inside browsers, phones and everyday tools, and those settings carry risks a corporate network does not. And how a protection layer is wired into the agent changes what it can see, so two products with the same headline score can be doing entirely different jobs.
A lock or good manners?
There is one design question underneath all of this that the industry has not settled. Should the attack be allowed to reach the model at all?
Most protection today works by letting the model see the hostile content and trusting it to behave. That should be the fallback, not the default. Once hostile content is inside the model's context, everything that follows is a judgement made by a system that has already been shown the attack. Nothing after that point is trustworthy in the way a lock is trustworthy.
When you test an attack against an agent, there are four possible outcomes. It never reached the model. It reached the model, and the model declined. Someone noticed and did not stop it. Nobody noticed. Most safety claims today do not distinguish between the first two, and yet they are the difference between a lock and good manners. A single number that covers both cannot tell you which one you have. The new methodology records them separately. It also separates noticing an attack from stopping it, and it treats blocking legitimate work as its own kind of failure, because a protection layer that gets in the way of the task is one users switch off. Published vendor claims get recorded too, and then marked as confirmed, unconfirmed or contradicted.
And there is the part nobody in this field likes to say out loud. These systems are not deterministic. They are getting better, quickly, and they will keep getting better. They will also keep giving you a slightly different answer to the same question. Would you accept a lock that works nine times out of ten? A protection layer for an AI agent runs thousands of times a day. The tenth time is not a rounding error; it is somebody's data.
From one internal thread to an industry standard
This did not need another vendor framework. It needed the vendors and the testers in the same room, agreeing on rules both sides could live with. That is what AMTSO has done since 2008, and it is why the work went there rather than into a Gen white paper.
Gen initiated it. It started in an internal thread about where the industry was heading, and one of us said the obvious thing: somebody independent should be measuring all of this. Including us. Under the AMTSO AI Security Working Group, with testing organizations and vendors from across the industry, the first version went from that conversation to an adopted, published standard in a couple of months. For a standards body, that is fast.
We have been building in this space since the start of the year: AI Agent Protection in Norton and Avast, Sage released as open source for the community, and AARTS and Skill IDs published as open standards anyone can implement. More is coming soon. And because a call for measurement carries little weight from someone unwilling to be measured, our intention is clear. When independent evaluations built on this standard start running, we want our own products on the list.
Measured, not announced
I do not want AI to be a thing people are afraid of. I also do not want it to be a thing people believe on faith. It should make life better, and for tens of millions of people it already does. What is missing is trust, and trust is not something a company can announce.
Being fearless about AI doesn’t mean trusting it blindly or ignoring its risks. It means understanding those risks, measuring them and building the protections that give people the confidence to use it.
Every previous era of computing learned this the same way: the claim comes first, the measurement comes later, and the gap between them is paid for by users. This time the measurement is arriving early. That is the whole point of it.
