OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red made the model its most robust release yet.

GPT-Red automates a type of safety evaluation for software systems known as red-teaming, which is typically done by a team of human testers. The aim is to find as many different ways to break or hijack a system as possible. The weak spots can then be patched before the final version of the software is released.

As LLMs become more complex and get used in a wider variety of tasks—especially in the form of agents, which can interact with computer files, websites, and third-party code as well as other agents—it’s hard for teams of people by themselves to keep up with all the types of attacks that might take place. “The risk surface grows and the blast radius also grows,” says Nikhil Kandpal, a research scientist at OpenAI who co-created GPT-Red.

OpenAI built GPT-Red to future-proof its safety testing process. “As more capable models become available, we will have already designed the system that can discover new modes of attack,” says Dylan Hunn, a research scientist at the company and fellow co-creator of GPT-Red. The researchers say it has already come up with new types of attack that had not been seen before.

OpenAI focused most of its efforts on a type of attack known as a prompt injection, where a hacker slips an LLM instructions to make it do things its developers or users do not want it to, such as copy confidential information, sabotage a company’s code base, or generate embarrassing or harmful output. In theory, such instructions can be hidden in any text that the LLM might encounter—in code or on a website, for example.    

Training dojo

To build GPT-Red, OpenAI’s researchers took an LLM that had not been trained as a hacker and set it up in what’s known as a self-play loop with several other models. Its goal was to try to attack the other models; their goal was to try to defend themselves. Over many rounds of play, GPT-Red became better and better at attacking other LLMs, and those LLMs became better and better at fending off the attacks.

The training took place in a kind of dojo that OpenAI had designed to mimic a range of scenarios in which LLMs might be deployed in the real world, including browsing the web, reading emails or calendar apps, and editing code.  

When GPT-Red found a new kind of attack, it would explore multiple different versions of it to find the most efficient one for specific scenarios. “Compared to a human red-teamer, the model is very, very good at finding exactly what will work, exactly what’s most effective,” says Hunn. “It’s extremely persistent about drilling down into an attack that it has discovered.”  

In particular, OpenAI claims that GPT-Red found a type of prompt injection attack that the researchers had not seen before, which they call a fake chain of thought. A chain of thought is a kind of diary in which an LLM makes notes to itself and keeps track of partial results as it works through problems. GPT-Red found a way to insert a fake entry into another model’s chain of thought that would trick that model into acting on spoofed information.

“It’s like if I told you that 1+1=3 and that you have verified this already,” says Chris Choquette-Choo, another research scientist on the team. “The model’s like, ‘Oh, okay, of course,’ and it just spits out 3.”

Jessica Ji, a senior research analyst who works on AI security at Georgetown University’s Center for Security and Emerging Technology (CSET), thinks the self-play loop that OpenAI used is a good approach. “The results look very promising,” she says.

OpenAI tested how good an attacker GPT-Red was by rerunning an experiment from 2025 in which human red-teamers tried to find weaknesses in an earlier version of GPT-5. When GPT-Red was set the same task, it was more successful at finding effective attacks than the humans had been.

OpenAI also tested GPT-Red against Vendy, a vending machine agent developed by Andon Labs, a company that assesses how well agents perform real-world tasks. GPT-Red was able to hack Vendy to make it change the prices of items on sale and cancel a customer’s order.

Defensive behavior

OpenAI says that when it tried out some of the strongest attacks that GPT-Red had come up with on its models, more than 90% of them worked against GPT-5 (released in August last year), and fewer than 23% worked against the new GPT-5.6.

GPT-Red isn’t perfect. It is not great at figuring out attacks that involve a back-and-forth conversation between hacker and target, something that human attackers would have few problems with. It is also not yet that great at using images, which can be used to pass text to models in prompt injection attacks.    

The company says that GPT-Red supplements the work of its human red-teamers. People can still find attacks it misses. One approach OpenAI is taking is to give GPT-Red an attack that humans came up with and ask it to find all the variations.

“I think human expertise will still be very important,” says CSET’s Ji. “It would be really useful to be able to distinguish where human testing is most needed.”

Unsurprisingly, OpenAI will not be releasing GPT-Red. The company is also confident that the super-hacker is stronger than any copycat model someone might try to create. The researchers say they have been working on the model for more than a year, backed by the compute resources of one of the richest companies in the world.

“It’s not a trivial thing that someone could easily do—you know, just go and train a super-attacker using this idea,” says Choquette-Choo.

Read more

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

PsiQuantum has a plan to make a massive quantum computer out of light

The machine that could change the world will be housed in a room that looks like a data center crossed with an ice cream factory. 

Inside, some 100 stainless-steel cabinets each hold hundreds of chips. On those chips, thousands of light particles will fly through a maze of optical switches and beam splitters. Each photon must be accounted for, because precisely measuring where it ends up will help answer questions that current computers might take millions of years to solve.

This computer, as described, does not exist. It’s the brainchild of a company called PsiQuantum, founded in 2016 by four physicists from UK universities. In a crowded field of deep-pocketed competitors with similarly fantastical visions, the company aims to be the first to build a useful quantum machine.

Read the full story on the company’s quest.

—James O’Donnell

MIT Technology Review Narrated: inside the world’s deepest and longest subsea road tunnel

—Niall Firth

I’m currently around 1,000 feet beneath the North Sea, in a dark, dank cave. It smells weird. And I’m increasingly aware of the pressure from millions of tons of seawater just above my head.

I’m under the iconic fjords of Norway to visit what will soon become the world’s longest and deepest subsea road tunnel—an exceptional engineering feat that will carry drivers deep beneath the North Sea.

I’m here to understand how you make a 16.6-mile highway that sits 1,280 feet below the sea at its deepest point. And also—at a time when it can feel hard to get anything done—to reassure myself that ambitious engineering is still possible. That we can still make things. 


This is our latest
story to be turned into an MIT Technology Review Narrated podcast, which we publish each week on Spotify and Apple Podcasts. Just navigate to MIT Technology Review Narrated on either platform, and follow us to get all our new content as it’s released.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 Meta allegedly used AI to target workers with health issues for layoffs
Their lawsuit says Meta relied on AI to create a termination list. (Guardian)
+ And pinpointed staff who took maternity or disability leave. (Reuters $)
+ One was allegedly informed the day before her water broke. (Ars Technica)
+ The layoffs aimed to offset Meta’s AI spending. (Gizmodo)
+ AI agents are not your “coworkers.” (MIT Technology Review)

2 OpenAI’s first consumer device will be a mobile smart speaker
The screenless device will serve as an “AI companion.” (Bloomberg $)
+ It’ll let you talk with ChatGPT. (Verge)
+ And use a camera and sensor to understand your environment. (Reuters $)
+ It’s set to launch next year. (Engadget)
 
3 The US military sent explosive drone boats into combat for the first time
They attacked an Iranian midget submarine and naval port. (Ars Technica)
+ Underwater drones may shape a war in Taiwan. (MIT Technology Review)
 
4 DeepMind’s CEO has called for a US-led body to test frontier AI models
Demis Hassabis wants the watchdog to vet national security threats. (FT $)
+ If dangers mount, it would coordinate an industry-wide slowdown. (Axios)
 
5 Data centers are set to add billions in power costs in 13 states
A power auction is slated to produce $6.3 billion in new charges. (NYT $)
+ Australia plans to govern the use of water and power for AI. (WSJ $)

6 xAI’s unpermitted power pollution hits Black communities hardest
Elon Musk’s xAI has been installing gas turbines without permits. (Reuters $)
+ We need to focus on Big Tech’s energy footprint. (MIT Technology Review)
 
7 Stripe and Advent have offered to buy PayPal for more than $53 billion
The payments giant and private equity firm have made a joint bid. (Reuters $)
+ Apple and Google Pay have eroded PayPal’s market share. (Bloomberg $)

8 DeepSeek plans to file for IPO as soon as this year
The Chinese AI pioneer is likely to list in Shanghai. (WSJ $)
+ Here’s why DeepSeek’s latest model matters. (MIT Technology Review)

9 A hard, lightweight “bio-metal” has been discovered in sea worm jaws
It could have applications in engineering. (New Scientist $)

10 A new $3,000 fitness suit electrocutes you to boost your gains
Celebrities love it—but not everyone’s a fan. (404 Media) 

Quote of the day

“By economic and engineering measures, generative AI might be the worst technology ever deployed.” 

—Alex Reisner, a staff writer at The Atlantic, explains why GenAI’s scaling problem is an engineering disaster.

One More Thing

""
FRANZISKA BARCZYK


Hackers made death threats against this security researcher. Big mistake.

In April 2024, an anonymous hacker began posting death threats on Telegram and Discord channels aimed at a cybersecurity researcher named Allison Nixon. It wasn’t long before others piled on. Someone shared AI-generated nudes of her.

They targeted Nixon because she had become a formidable threat. As chief research officer at the cyber investigations firm Unit 221B, named after Sherlock Holmes’s apartment, she had built a career tracking cybercriminals and helping get them arrested. 

For years, Nixon had lurked quietly in online chat channels or used pseudonyms to engage with perpetrators and bring them to justice. Now, she resolved to unmask the people behind the death threats—and take them down for crimes they admitted to committing. 

Find out why they learned to regret their choice of target.

—Kim Zetter

We can still have nice things

A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.)

+ A musician has discovered the true masters of metal breakdowns: birds.
+ Photographer Fontanesi’s surreal photo splits transform everyday images into spectacular hybrid scenes.
+ Over 30 actors, filmmakers, and friends recount how Steven Spielberg infiltrated Hollywood in this terrific article.
+ Who would win the World Cup if less important things than soccer decided it, like life expectancy and happiness? A new game tests your knowledge.

Read more
1 … 96 97 98 99 100 … 3,357