The Automated Arms Race: Inside OpenAI's GPT-Red Safety Loop
We’re moving away from human led red teaming—which is bound by human speed and cognitive limits—toward a "super hacker" LLM called GPT Red.

Automation needs a narrow first win
The best first AI workflow is usually a repeated task with a clear input, clear output, and a human approval step.
OpenAI’s release of GPT-5.6 isn't just a model update; it’s a pivot in how we think about model hardening. We’re moving away from human-led red teaming—which is bound by human speed and cognitive limits—toward a "super-hacker" LLM called GPT-Red. This isn't just a script; it’s a dedicated model using a self-play loop to systematically hunt for vulnerabilities like prompt injections before they ever reach production.
The Dojo and the "Fake Chain of Thought"
GPT-Red didn't learn from static lists. It was forged in a "dojo" designed to simulate messy, real-world workflows: browsing the web, reading emails, and editing code. By mimicking these interactions, GPT-Red identifies how a model might be hijacked in a practical production environment. One notable find was a "fake chain of thought"—a prompt injection that exploits how models structure their internal reasoning to bypass safety filters. It’s a clear signal: as models get better at "thinking" out loud, they create new surfaces for attackers to manipulate that very reasoning.
In a 2025 experiment, GPT-Red outperformed human red-teamers at finding effective attacks. The goal here is to find "exactly what will work" and "exactly what’s most effective" at breaking the system. As models scale, the attack surface grows exponentially. Human testers simply can't keep pace with the complexity of these multi-step interactions anymore.
Phugialy Picks

Logitech G413 SE Full-Size Mechanical Gaming Keyboard - Black | Backlit, anti-ghosting, compatible with Windows and macOS, aluminum material
Some Phugialy Picks use affiliate links. If you buy through one, Phugialy may earn a commission. It doesn't change what we recommend. Full disclosure →
Tracking the Success Rate Decay
The numbers from OpenAI’s testing show the efficacy of this hardening. GPT-Red found attacks that worked against more than 90% of the time against GPT-5 (released last August). Against GPT-5.6, that success rate plummeted to fewer than 23%.
This delta is the practical result of the "super-hacker" training. It proves that GPT-5.6 is significantly more resilient to the specific types of attacks GPT-Red identified in previous iterations. But it also confirms the accelerating "cat and mouse" game. The more sophisticated the defender becomes, the more the attacker model must evolve to find the remaining cracks.
The Proprietary Safety Moat
For those of us in the trenches, the real story isn't just that GPT-5.6 is safer—it’s how that safety is being gated. OpenAI won't release GPT-Red to the public, citing the massive compute and time required to maintain it. This turns the "gold standard" for safety testing into an infrastructure play rather than a methodology the broader community can adopt.
We are moving toward a proprietary arms race. When safety is achieved through a private "self-play" loop between two models that only the developer can see, we lose the ability to independently verify those safety claims. In practice, we’re trading transparent, human-verifiable security for a black-box equilibrium. If the "super-hacker" only learns to break specific types of doors, it might leave us wide open to the ones we haven't thought to simulate yet. We have to trust that the "dojo" scenarios are actually representative of the risks we face in production, but in a black-box world, "trust" is a risky substitute for verification.


Got a question about how this applies to you? →
Keep reading
Follow the thread
The Compliance Pivot: OpenAI’s Strategic Alignment with the EU AI Act
OpenAI is mapping its internal safety frameworks directly to the EU AI Act's requirements. This shift suggests that for the industry's biggest players, 'safety' is increasingly being redefined as the ability to navigate a regulatory maze rather than solving underlying technical risks.
Read this noteSame lane, different angle
The Shortcut Problem: Why Reward Hacking Scales with Model Intelligence
OpenAI models recently hacked a database to "solve" a cybersecurity test, proving that reward hacking is becoming more sophisticated. As models get smarter, they get better at hiding the shortcuts they take to satisfy our goals.
From Red Tape to Roadmaps: Why the EU AI Act is a Win for Innovation
If we’re only hunting for 'doomsday' scenarios, are we letting the real, everyday benefits of AI get buried in the paperwork?