OpenAI has introduced GPT-Red, an automated AI system designed to find security vulnerabilities in its language models.
GPT-Red takes its name from cybersecurity red teaming, which is the practice of deliberately attempting to break a system to identify weaknesses before attackers can exploit them.
“GPT‑Red learns through adversarial self-play, where its goal is to prompt inject a variety of challenging defender models,” OpenAI wrote. “Every successful attack that GPT-Red finds is used to improve these defenders, pushing GPT‑Red to continuously find broader and more complex failures.”
In one case study, OpenAI said the system manipulated an autonomous vending machine agent into lowering prices, ordering discounted inventory, and canceling another customer's order before the vulnerabilities were disclosed and addressed.
GPT-Red follows years of cybersecurity efforts by OpenAI after the public launch of ChatGPT.
OpenAI's announcement reflects a broader shift toward using AI to secure AI.
According to OpenAI, GPT-Red will remain an internal tool because it contains intentionally developed offensive capabilities.
“We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today's models can be used to make tomorrow's models more robust, aligned, and trustworthy,” they said.


















