August 26, 2026|4 min read

Hiring a Red Team Isn't Assurance. Owning the Verdict Is.

A vendor red-teamed OpenAI, Anthropic and Meta, and it went off the rails. Third-party AI testing isn't a control until someone owns the verdict.

Written by Carlos Alvidrez, with AI assistance in research · How we use AI

Hiring a Red Team Isn't Assurance. Owning the Verdict Is.

Photo by Preillumination SeTh on Unsplash

A small Israeli firm called Irregular landed the job every security vendor dreams about: assess how secure the frontier models from OpenAI, Anthropic and Meta really were. Then, as The New York Times reported, it made a mistake, and the tests went off the rails.

The reflex is to file this under competence. A vendor slipped. Tighten the process, move on.

That reading is comfortable and wrong. This is a governance story. Nobody had written down, before the first prompt, who owned the scope of the test, where the failure line sat, and who held the authority to stop. When that contract is missing, third-party evaluation is not a control. It is a performance with a slide at the end.

"You are not governing. You are negotiating with your own excitement."

A test is not a verdict

Here is the sleight of hand that reassures everyone and protects no one.

A lab commissions an outside evaluation. The report comes back. The lab points to it: "independently tested." The board relaxes. The press release writes itself.

But a test only produces observations. Ten thousand adversarial prompts, a pile of transcripts, a score. None of that is a decision. Somebody still has to convert the observations into one word the whole company will act on: pass, or fail.

That conversion is the entire control. And in most of these arrangements, nobody owns it.

So when the score is ambiguous, or the vendor makes a mistake, or a result lands in the gray zone nobody scoped, there is no one whose job is to say "this does not ship." The observations pile up. The launch calendar keeps moving. The gap between "we tested it" and "we decided it was safe" is where the incident lives, and no one is standing in it.

Off the rails is not what a bad test looks like. It is what an unowned test looks like the moment reality stops cooperating.

The missing piece is a written contract of who decides

Not a legal contract. A governance one. One page, agreed before anyone runs a single prompt, that assigns three things to named humans:

  • Scope. What is in the test and what is out, written down and signed off before the work starts, so the lab and the vendor are not quietly assuming two different boundaries.
  • Failure conditions. The exact line that separates a pass from a fail, defined before the results arrive. Set the bar after you see the score and you will always find a reason the score clears it.
  • Authority to stop. One named person whose "fail" halts the release. Not advises. Halts. If their verdict is a recommendation the launch team can overrule, it was never a control.

Notice what these have in common. Each one has to exist in writing, owned by a specific person, before the test runs. Decide any of them afterward, in the heat of a shipping deadline, with a hundred-million-user launch on the line, and you are not governing. You are negotiating with your own excitement.

The frontier labs did not run into trouble because Irregular is bad at testing. They ran into trouble because "assurance" was assumed rather than assigned. Everyone believed someone else owned the verdict. No one did.

Your Monday move is small and uncomfortable. Pull the last external assessment your team paid for. Find the three sentences that say who owned the scope, who set the failure line, and who could have stopped the release. If those sentences do not exist, you did not buy a control. You bought a story you told yourself, with a logo on it.

Then write the one page for the next one. Get it signed by the person who can actually halt a launch. If you cannot name that person in a single breath, that is the finding. Everything else is theater.

So, the gut-check. The next time a clean report from an outside tester lands on your desk, ask the only question that counts before you read a word of it: if this thing should not ship, who here has the authority to say so, and does everyone already agree it is binding?

If the answer only shows up after the results do, you already know what you bought.

Sources

Related governance guides