# Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

> OpenAI has built GPT-Red, a model trained not to be helpful but to break other models. Pointed at its own systems, it drove the rate of successful attacks on GPT-5.6 down from more than 90 percent to under 23 percent, and it did so by outperforming the human red-teamers running the same tests. Along the way it surfaced a new attack the company calls fake chain of thought, where false statements are slipped into a model's reasoning steps to steer its conclusions. The idea of automating adversarial testing is not new, but the results here are a reminder that the same capabilities that make models dangerous also make them useful defenders. An AI that can find holes faster than people can is a powerful safety tool. Notably, OpenAI has no plans to release GPT-Red, treating it as a proprietary edge rather than shared safety infrastructure. That choice is its own story: the best AI security tooling may end up locked inside the labs that need it least, while everyone else waits.

_Section: [Daily AI Updates](https://www.wortins.com/daily-ai) · Source: MIT Technology Review · Published Sunday, July 19, 2026_

## Wortins' read

OpenAI has built GPT-Red, a model trained not to be helpful but to break other models. Pointed at its own systems, it drove the rate of successful attacks on GPT-5.6 down from more than 90 percent to under 23 percent, and it did so by outperforming the human red-teamers running the same tests. Along the way it surfaced a new attack the company calls fake chain of thought, where false statements are slipped into a model's reasoning steps to steer its conclusions. The idea of automating adversarial testing is not new, but the results here are a reminder that the same capabilities that make models dangerous also make them useful defenders. An AI that can find holes faster than people can is a powerful safety tool. Notably, OpenAI has no plans to release GPT-Red, treating it as a proprietary edge rather than shared safety infrastructure. That choice is its own story: the best AI security tooling may end up locked inside the labs that need it least, while everyone else waits.

## Source

[Read the full story at MIT Technology Review](https://www.technologyreview.com/2026/07/15/1140514/meet-gpt-red-an-llm-super-hacker-openai-built-to-make-its-models-safer/)

## Related coverage

- [Healthcare AI Diagnostics Adoption Accelerates Despite Safety Concerns](https://www.wortins.com/story/healthcare-ai-diagnostics-adoption-accelerates-despite-safet-5b905590) — [Radiology Business](https://radiologybusiness.com/topics/artificial-intelligence/navigating-ai-diagnostic-dilemma-healthcares-no-1-patient-safety-concern-2026)
- [Moonshot AI Unveils Kimi K3, Closes Gap With US AI Rivals](https://www.wortins.com/story/moonshot-ai-unveils-kimi-k3-closes-gap-with-us-ai-rivals-448574b5) — [Fortune](https://fortune.com/2026/07/16/moonshots-kimi-k3-pushes-chinese-ai-into-fable-level-territory/)
- [Anthropic and Blackstone Bet on AI Implementation, Not Models, as Trillion-Dollar Business](https://www.wortins.com/story/anthropic-and-blackstone-bet-on-ai-implementation-not-models-d5249b46) — [TechCrunch](https://techcrunch.com/2026/07/15/anthropic-blackstone-bet-the-next-trillion-dollar-ai-business-is-implementation-not-models/)
- [IBM Stock Plummets 25% on Earnings Miss and Spending Shift to Hardware](https://www.wortins.com/story/ibm-stock-plummets-25-on-earnings-miss-and-spending-shift-to-bd21cf56) — [CNBC](https://www.cnbc.com/2026/07/14/ibm-warns-second-quarter-earnings-fell-short-of-expectations.html)
- [Anthropic Prepares for October IPO as Fable 5 Dominance Solidifies](https://www.wortins.com/story/anthropic-prepares-for-october-ipo-as-fable-5-dominance-soli-98c09b96) — [GuruFocus](https://www.gurufocus.com/news/8960658/anthropic-prepares-for-potential-ipo-outpacing-openai)
- [Stanford Letter: 200+ Economists Warn of AI Disruption Larger Than Industrial Revolution](https://www.wortins.com/story/stanford-letter-200-economists-warn-of-ai-disruption-larger--61b57833) — [Intellectia](https://intellectia.ai/blog/ai-infrastructure-investment-july-2026)

---

_Curated and written by [Wortins](https://www.wortins.com) — The daily AI briefing. Every story links to its original source; the "Wortins read" on each is our own original analysis. [About Wortins & our editorial approach](https://www.wortins.com/about)._
