OpenAI presented on Wednesday July 15, 2026 an automated red-teaming tool named GPT-Red, responsible for strengthening the resistance of GPT-5.6 to prompt injection attacks. The concept is based on a simple observation: human penetration testing methods no longer keep pace with the capabilities of the models. The challenge is now growing, as these flaws directly affect the security of autonomous agents.

In brief
- GPT-Red passed 84% of internal prompt injection evaluation scenarios, compared to 13% for human red teams.
- OpenAI trained GPT-Red using self-adversarial reinforcement learning to harden GPT-5.6 before deployment.
- The Ethereum Foundation has also deployed AI agents to audit its critical network infrastructure in July 2026.
An AI model that attacks itself to protect GPT-5.6
The idea of using an AI to boost another AI is not new, but OpenAI is pushing it here to an operational stage. The group has already paved the way by deploying AI agents to track critical flaws in its own network, a strategy that is also found on the Ethereum side, where the foundation has entrusted autonomous agents with the red-teaming of its infrastructure.
GPT-Red is rooted in this automated offensive security logic. GPT-Red takes its name from “red teaming”, this cybersecurity practice which consists of deliberately attempting to break a system to identify its weaknesses before an attacker exploits them.
OpenAI explains that the model was trained using self-play reinforcement learning. It generates increasingly sophisticated prompt injection attacks, while defending models learn to resist them. Each successful assault then feeds GPT-5.6 trainingwhich emerged more robust even before its deployment.
In one case study cited by OpenAI, the system manipulated an autonomous agent operating a vending machine, tricking it into lowering prices, ordering discounted inventory, and canceling another customer’s order.
The flaw was reported and corrected before any real exploitation. The example shows how a prompt injection can transform an assistant into a hijacked tool, without the user realizing it.
An 84% score that crushes human red teams
The number that stands out in OpenAI’s announcement is the difference in performance measured internally. In the same evaluation scenarios, GPT-Red succeeded in 84% of prompt injection attacks, compared to only 13% for human red teams.
OpenAI justifies this automation in a message published on As model capabilities increase, safety and alignment must evolve as well », writes the company.
Red-teaming is essential, but current approaches are difficult to scale, creating a critical bottleneck. GPT-Red is one way we solve it.
The model works by adversarial self-confrontation, specifies OpenAI. “ GPT-Red learns through adversarial self-confrontation, its goal being to inject prompts into a variety of difficult defender models », explains the company.
Each successful attack that GPT-Red discovers serves to improve these defenders, pushing GPT-Red to continually find larger, more complex failures.
The loop feeds on itself, and that’s precisely what the researchers were aiming for: an engine for continuous improvement rather than a one-off testing campaign.
From ChatGPT to automated red-teaming, security goes to scale
GPT-Red continues several years of cybersecurity efforts launched by OpenAI following the public success of ChatGPT. The company created its OpenAI Red Teaming Network in 2023, recruiting external researchers to probe its models for flaws before publication.
The move to the automated model marks a shift in gear, since AI produces attacks on a scale unattainable for humans alone.
This announcement is part of a broader movement: that of AI securing AI. Earlier in July 2026, the Ethereum Foundation reported that it deployed AI agents to audit its critical network infrastructure, discovering a vulnerability in software used by its consensus clients.
Researchers have noted that AI agents explore larger code bases than humans, but the real challenge has slipped: it’s no longer about spotting bugs, but proving which ones are actually exploitable.
OpenAI keeps GPT-Red under lock and key, but sees it as a virtuous circle
OpenAI keeps GPT-Red as a purely internal tool. The model contains offensive capabilities developed voluntarily, which excludes any public dissemination. However, the company sees this as the start of a virtuous circle.
“ We believe with GPT-Red we have begun to unlock a similar ripple effect for security, where today’s models serve to make tomorrow’s models more robust, aligned and trustworthy. », she concludes.
The challenge now is to transform this internal lead into lasting confidence among regulators and users.
In short, OpenAI has made the automated attack a shield for GPT-5.6, with a performance gap that commands respect: 84% success for GPT-Red compared to 13% for humans. This shift towards an AI that secures another AI is reshaping the security posture of the industry, from laboratories to blockchains.
Maximize your Tremplin.io experience with our ‘Read to Earn’ program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.
