17 Jul 2026
Inherited robustness: you cannot run the lab's test, and the evidence expires (GPT-Red). Model choice as a security control. News-pegged.
Most organisations believe they have a rough sense of how exposed their AI systems are.
Something OpenAI published this week shows how much of that judgement rests on evidence they cannot reproduce.
It has built GPT-Red, an internal automated attacker trained specifically to break AI systems. By the company's own account, it succeeded against GPT-5.1 in 84% of cases on one set of novel prompt-injection scenarios, compared with 13% for human red-teamers. It also broke a live vending-machine agent in OpenAI's office, changing prices and cancelling another customer's order.
GPT-Red is being kept internal, and every frontier lab runs red-teaming work it does not release. That is defensible: an instrument built to discover attacks is also useful to attackers.
But it creates an asymmetry. Customers can test their applications, permissions and surrounding controls. What they cannot reproduce is the lab's strongest test of the model itself.
That layer of robustness is inherited. The lab attacks, hardens and ships the model. You receive the result and decide whether the evidence is good enough.
The evidence also expires. One attack class GPT-Red discovered succeeded more than 95% of the time against GPT-5.1 and less than 10% against GPT-5.6 Sol. Model choice is therefore not merely a question of cost and capability. It is a security control.
Plenty of organisations are still running older models for sensible reasons: an integration nobody has recertified, a vendor that has not shipped, or procurement that treats standing still as prudence. Each is also a security position, often taken by people who cannot independently size the model-level risk.
A newer model does not automatically make a system secure. The tools, permissions, harness and execution environment still determine what an agent can do when it fails. But one important layer can change sharply when the model changes, and the best measurement of that layer remains with the vendor.
So the useful question is not whether your AI is secure.
It is when you last checked which model your systems are actually running, what it is allowed to do, and who decided that risk was acceptable.
The question is, do you think or do you know your model is secure?
Originally published on LinkedIn
Want to apply this to your business?
If this sparked a useful question, let’s talk about where AI, automation, or product strategy can create practical leverage in your organisation.
Start the conversation