This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

Reading OpenAI’s account last week of how some of its models broke their containment and hacked into the computer systems of Hugging Face, another AI company, was the first time I got genuine chills about what large language models are now able to do. But this is a case of human hubris, not rogue AI.

I am not an alarmist. In fact, I have been pushing back against AI scare stories for years. Even so, this incident crossed a line. I think it’s the clearest illustration yet of how the people building and testing this technology do not fully understand what they’re doing. OpenAI could—and should—have seen this coming.

Here’s what happened, at least according to the two companies involved. A couple of weeks ago, OpenAI started testing the hacking abilities of some of its new models, including GPT‑5.6 Sol (released in June) and what OpenAI describes as “an even more capable pre-release model.”

OpenAI pitted its models against a benchmark called ExploitGym, released in May, which challenges LLMs to find ways to exploit real-world vulnerabilities found in commonly used software.

To see what they could do, the researchers removed most of their cybersecurity guardrails. Then they ran the models inside a sandbox that was cut off from the internet except for one link to a third-party piece of software that acted as a proxy to the outside world, and let them install code that they needed to beat ExploitGym.

On July 9, according to reporting by Reuters, OpenAI’s models started trying to break through the proxy. They found an unknown bug in the proxy’s software and used it to access the internet. From there, they broke into Hugging Face’s computer systems on July 11, apparently looking for data sets and solutions that would help them complete their task. Hugging Face announced the hack on July 16. 

OpenAI did not realize (or at least did not reveal) that its models were involved until July 21, around 10 days after they broke containment and a week after Hugging Face had shut down the attack and alerted the FBI.

In a statement given to MIT Technology Review, OpenAI says: “We are conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we will publish a technical report of our learnings for everyone.” The firm also confirmed that its researchers were properly using existing safety guidelines and procedures at the time.

Wake-up call

OpenAI has said the event was unprecedented—and in many ways it was. This was the first time outside of a simulation that LLMs escaped what was thought to be a secure sandbox, accessed the open internet, and attacked an unrelated organization. It’s a wake-up call that shows just how good the latest LLMs are at finding and exploiting vulnerabilities in real-world software with little or no human guidance.

And yet at the same time, what OpenAI’s models did is something this technology has done for years. Give a model a goal and it will very often achieve that goal in unexpected ways, finding loopholes that look like cheats. OpenAI itself has studied this behavior.

A decade ago, it shared results of an experiment in which a model was tasked with beating a video game called CoastRunners. Human players take it for granted that the way to do this is by racing a boat through a series of flags to the finish line, racking up points for each flag you hit. OpenAI’s model figured out that you could get a high score by spinning in a circle and hitting the same three flags over and over again. There have been dozens of similar examples from researchers since. AI will always find a way.

“Despite repeatedly catching on fire, crashing into other boats, and going the wrong way on the track, our agent manages to achieve a higher score using this strategy than is possible by completing the course in the normal way,” OpenAI wrote in a blog post about the CoastRunners experiment in 2016. “While harmless and amusing in the context of a video game, this kind of behavior points to a more general issue … it is often difficult or infeasible to capture exactly what we want an agent to do.”

I couldn’t help thinking about CoastRunners when I read OpenAI’s blog post about the Hugging Face attack: “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal … After gaining internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.”

Last week’s news was not about rogue AI, despite the headlines. It was about models achieving the goal they had been given: Find ways to exploit vulnerabilities in software. The fact that those models then behaved in a way OpenAI had not anticipated isn’t surprising. But it is worrying.

Back in 2016, OpenAI had this to say about its CoastRunners bot: “More broadly it contravenes the basic engineering principle that systems should be reliable and predictable.” A decade on, those basic engineering principles are still AWOL.  

Read more

Outside the small town of Paducah, Kentucky, a wealth of uranium is locked away in thousands of storage cylinders filled with waste material from a now-closed nuclear enrichment facility. Lasers could help get it out.

A company called Global Laser Enrichment (GLE) is looking to reprocess this old material with a new technology called laser enrichment. It could be more efficient than conventional enrichment methods, allowing the company to refresh the material and produce feedstock at the same concentration as a natural mined source. And in the future, the company claims, laser enrichment could be used to make material for nuclear fuel, including the kind used in advanced reactors.

Nuclear power provides about 9% of global electricity today, and that fraction could tick up as major world powers like the US and China look to build new reactors, including some based on next-generation technology. New, cheaper methods to obtain fuel could help ensure that those nuclear projects stay on track.

Naturally occurring uranium is largely made up of uranium-238 (over 99%) and uranium-235 (about 0.7%). Uranium-235 is the fissile type, meaning that, when hit with slow low-energy neutrons, it can sustain a chain reaction that generates electricity. So reactors generally use material with a higher concentration of U-235 than what’s pulled from the ground. Today’s conventional reactors usually use low-enriched uranium, typically is about 5% U-235, though some advanced reactor designs will use fuel that’s up to 20% U-235.

Today, centrifuges are the dominant tech used to enrich uranium. The equipment essentially takes uranium-containing material and spins it around incredibly quickly, so the heavier material (which contains U-238) spins out to the edge, while the lighter material (which has U-235) stays closer to the center. (If you’ve ever swung a mustard bottle to get the last of it out, you’ve used the same basic idea behind a centrifuge.) Then the material that has a higher concentration of U-235 can go on to be made into nuclear fuel.

Laser enrichment, on the other hand, takes advantage of the fact that all molecules vibrate and rotate at an atomic scale in ways that depend on their specific material. Even different uranium isotopes have distinct fingerprints.

Lasers are so precise they can target one particular material (like molecules that contain U-235, for example). If you shine a laser at a mixture, you can selectively excite just the material you’re targeting, giving it a bit more energy. This changes the way it behaves, which can make it easier to separate out the material you want using chemical or physical methods.

A wide range of separation approaches have been developed in research and industry. Some aim to electrically charge U-235 atoms, allowing them to be moved with electrostatic or magnetic fields. Others change how the material reacts chemically. 

The details of GLE’s specific technology are classified, and company officials declined to share how the process works. 

There’s been interest in using lasers for uranium enrichment for decades, says Charles Forsberg, a principal research scientist in nuclear science and engineering at MIT. 

However, in their early days lasers tended to be high-maintenance, unstable and difficult to operate. They’ve improved dramatically, making laser enrichment a more attractive prospect than it was during the early research.

Even more than technological improvements, a recent geopolitical shift could boost new enrichment technology. Russia has the largest uranium enrichment ecosystem in the world, and the country has historically dominated the market. “Nobody in the West was going to build a new enrichment plant while the Russians flooded the world with enriched uranium,” says Forsberg. 

Since the start of the Ukraine war, however, countries including the US and UK have taken steps to limit or ban imports of Russian uranium. That’s opened the door for companies to set up new enrichment operations, including some that use new technologies, Forsberg says.

Demand for fuel is increasing as countries look beyond Russia for uranium supply. “The gap is just becoming bigger and bigger, and this technology is right in the middle,” says Christo Liebenberg, president of LIS Technologies, one of the companies aiming to build laser enrichment capacity in the US.

LIS Technologies was founded in 2023, and the company recently purchased a 200-acre site in Oak Ridge, Tennessee. It’s currently in the pre-application process with the US Nuclear Regulatory Commission for its facility. The company plans to take in natural-grade uranium and make a product that’s roughly 5% U-235, though it hopes to eventually make more concentrated material that can be used as fuel for next-generation reactors.

GLE is taking a different approach: Rather than using its technology to enrich freshly mined material to the 5% concentration that can be used in fuels, it’s hoping to start by rehabilitating old waste.

The company has a contract with the US Department of Energy to reprocess waste material at the enrichment site in Paducah. The facility could enrich up to 200,000 metric tons of material that contains small amounts of uranium leftover from an older enrichment process.

GLE is taking the material that’s at least 0.25% U-235 and enriching it to about 0.7%. That material can then be further processed and slotted into the uranium supply chain in place of freshly mined material. “It’s kind of like a large aboveground uranium mine for us,” says Nima Ashkeboussi, vice president of government relations and communications at GLE.

While each one of its units is more complex and expensive than a centrifuge, far fewer are needed to do the same work. A similar centrifugation plant would have many thousands of centrifuges working together, but a full-scale plant using GLE’s laser enrichment process would have fewer than a thousand of its units, says Stephen Long, the company’s CEO. Up-front investment should be smaller, and operating costs are also expected to be lower, partly because the process uses less energy than centrifuges, Long says.

GLE has a testing facility in Wilmington, North Carolina. In fall 2025, the company completed a demonstration pilot, processing several hundred kilograms of uranium. It decommissioned that system and is currently putting together a new demonstration at the North Carolina plant, which would show how the technology works at commercial scale.

The company also applied for a license with the US Nuclear Regulatory Commission for its proposed facility in Paducah. The final safety evaluation should be finished in November, and the final approval should come in 2027, Long says. The plan is to start processing material at the plant by 2030.

In the long run, there’s plenty of uranium on the planet to keep reactors running for decades. But as interest in nuclear power grows and the geopolitics of fuel shift, there could be short-term gaps or price spikes that alternative sources could help smooth out.

Laser enrichment plants could turn out to be cheaper than existing technologies, says Stephen Greene, a senior fellow at the Nuclear Innovation Alliance. But as with most new technologies, “you don’t really know until you try to build one.”

Read more

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

How lasers could help provide fuel for nuclear reactors 

Nuclear power provides about 9% of global electricity today, and that fraction could tick up as countries look to build new reactors. New, cheaper methods to obtain fuel could help ensure that those nuclear projects stay on track.

One of those methods is called laser enrichment. It allows you to separate out the material you want (in this case, uranium) from others in a mixture of old waste.

A company called Global Laser Enrichment (GLE) is about to start testing whether the technology works at commercial scale. Read our story about their efforts.

—Casey Crownhart

The quest to keep organs alive outside the body

It’s super difficult to freeze organs. Once ice forms in them, they’re done. The ice crystals create all kinds of damage and render the organs unusable. That hasn’t stopped many researchers from trying.

In new research, one team has been able to supercool the kidneys of pigs and preserve them for days. The kidneys survived being stored at −4 °C (25 °F) and eventually reimplanted back into pigs. And that’s just the latest development in a field that is positively buzzing.

Read about why it’s such an exciting time for organ preservation—and what could be coming next. 

—Jessica Hamzelou

This story is from The Checkup, our weekly biotech newsletter. Sign upto receive it in your inbox every Thursday.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 Silicon Valley is split over how to respond to Chinese AI
It boils down to whether AI models should be open or closed. (NYT $)
+ Nvidia, Microsoft and Meta warn that restricting open models would backfire. (CNBC)
+ AI companies are spending record sums on lobbying Washington. (FT $)
+ China’s AI models have Trump’s AI world at war with itself. (MIT Technology Review)

2 Trump can’t post his way out of this war 
Iran has revealed hard limits to his ability to bend reality to his will. (Atlantic $)
+ Trump has been forced to abandon further escalation due to dwindling munitions stockpiles. (NYT $)
+ An Iranian strike on CIA facilities has raised questions about Russian involvement. (Reuters $)

3 Wildfires are surging across Europe
Repeated heat waves have turned parts of the continent into a tinder box. (BBC)
+ One of the fires forced NASA to evacuate a tracking station in Spain. (Ars Technica)
+ Americans are increasingly grappling with smoky skies too. (Atlantic $)

4 OpenAI didn’t notice its agent going on a days-long hacking spree
It only cottoned on after the threat was contained and the FBI had been alerted, sources say. (Reuters $)

5 The AI jobs wipeout still hasn’t arrived
In fact, a lot of companies are now embarking on hiring sprees. (WSJ $)
+ AI’s impact is increasingly falling short of expectations. (The Guardian) 
+ Here’s a much-needed reality check on the AI jobs hysteria. (MIT Technology Review)

6 A six-year-old girl died in a Chinese gene-editing trial
Experts say it should have never been allowed to go ahead. (Science)
+ This baby boy was treated with the first personalized gene-editing drug. (MIT Technology Review)

7 What it’s like to use a North Korean smartphone
They’re growing in popularity—but represent another avenue for government control. (WP $)

8 The FCC’s ban on foreign-made drones is not working
You can’t change global supply chains at the stroke of a pen. (The Verge $)

9 The “summer of Ludd” shows it’s fun to be a Luddite 
A growing anti-tech movement is all about raw, anarchic joy. (404 Media)
+ We’re in the era of AI malaise. (MIT Technology Review)

10 Why Jimothy the racoon is the internet’s latest obsession 🦝
It’s his irresistible combination of chaos and cuteness. (BBC)

Quote of the day

“I think that [AI] should stand for artificial idiot.”

—Marian Agnew, a nine-year-old, from Norman, Oklahoma, tells Wired she’s not impressed by AI models’ tendency to make up facts.  

One More Thing

three silhouetted people in a boat crossing the water in the dark toward a beam of light
KATHERINE LAM

Inside a romance scam compound—and how people get tricked into being there  

Gavesh’s journey started, seemingly innocently, with a job ad on Facebook promising work he desperately needed. 

Instead, he found himself trafficked into a business commonly known as “pig butchering”—a form of fraud in which scammers form close relationships with targets online and extract money from them. 

The Chinese crime syndicates behind the scams have netted billions of dollars, and they have used violence and coercion to force their workers to carry out the frauds from large compounds, several of which operate openly in the quasi-lawless borderlands of Myanmar. 

Read our story about these scam syndicates and how they could be broken up. 

— Peter Guest and Emily Fishbein

We can still have nice things

A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.)

+ I want to make every single one of these delicious-looking Korean dishes. 
+ If you like origami, you’ll love this guy’s tutorials.
+ How to deal with those old gadgets that are collecting dust in your drawer.
+ Enjoy these old art deco public transport posters from London.

Read more

Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy.

Each is an expert in its domain. But they all have their own distinct knowledge and objectives. Today they can exchange data, but they are not yet able to actually coordinate patient care without a human making the decisions.

“The intelligence is already there. What is missing is the connective tissue that turns four strangers into one team,” explains Vijoy Pandey, senior vice president and general manager of Outshift by Cisco.

This “connective tissue” comes from adding a semantic layer—what Outshift calls the “Internet of Cognition”—that enables agents across domains to work together and, critically, “think” together through shared intent, context, and reasoning.

This semantic layer relies on a connectivity layer beneath it called the “Internet of Agents,” which allows autonomous agents to discover one another, prove identity, and exchange messages across domains.

When used together, they enable “the next step on the road to distributed artificial superintelligence,” says Pandey.

From solo silicon savants to the ‘Internet of Cognition’

For years, the AI industry has been focused on growth. Scaling vertically has led to bigger models, trained on more data with more compute. This has produced the reasoning capabilities that can be like a “brain” for AI agents, which can perceive, reason, and act in digital environments.

While vertical scaling can produce more capable agents perpetually, to enable agentic problem solving across different systems, companies, and platforms the next axis of scale must be horizontal, says Pandey.

Multi-agent systems are already being explored in areas like software engineering, drug discovery, and scientific simulations, but their performances so far have been underwhelming. One study finds a failure rate of between 41% and around 87% when evaluating seven open-source multi-agent systems.

“Connected agents handle coordinated action well; taking a task whose shape they have seen, divided and passed around,” Pandey explains. “What they cannot do is hold a goal in common and reason toward something none of them was trained to solve.”

“The gap is architectural, not a prompting problem,” Pandey adds. “Without the right coordination layer, naive multi-agent setups can perform worse than a single agent. The step change is that team of agents converging on its own, on a new problem, with no human stitching the seams.”

To reach this goal, Pandey says Outshift has built a connectivity layer called AGNTCY, an open-source project now under the Linux Foundation. AGNTCY allows agents across different systems, companies, and platforms to find each other, prove identity, and exchange messages through open, standardized protocols.

And, as Pandey explains, this allows the Internet of Cognition thesis to take a step further. It creates a semantic layer that allows agents to align goals (share intent), pool institutional knowledge and compound memory (share context), and make collective trade-offs (share reasoning).

Pandey likens this progression to that of humans: “For hundreds of thousands of years humans got individually smarter, and the gains died with each person who made them,” he explains. “Around 70,000 years ago that changed, when humans learned to share intent, build cumulative knowledge, and reason collectively. That is when scattered individuals became civilization.

“Agents are at the same threshold. We have built the silicon geniuses and given them agency. What they lack is the layer that let humans go collective,” he says.

First steps to distributed superintelligence

Enabling agents to work collectively rests on three pillars in the tech stack:

Shared intent through cognition state protocols: Cognition state protocols are the semantic handshake that allow agents to agree on a goal before they act and then negotiate toward it. Outshift has created an open-source coordination layer called Mycelium, which organizations can clone and use against their own agents.

“We found that unstructured groups reached a decision about a third of the time across 14 scenarios,” says Pandey, speaking about internal testing. “A coordination protocol that makes agents declare a goal, surface missing information, and resolve conflicts before acting raised that to 93%.”

Shared context through cognition fabric: A cognition fabric is a shared institutional memory and communication mesh that allows agent insight to compound over time rather than resetting each session. This policy-governed context layer solves the problem of “organizational amnesia,” says Pandey, by ensuring the baseline intelligence of the systems only ever goes up.

Shared reasoning through cognitive amplifiers and guardrail technologies: Two kinds of cognition engine can be used together to enable shared reasoning. Cognitive amplifiers speed up shared reasoning and modeling, and guardrail technologies (GATs) create security, cost, and compliance frameworks. Humans are active contributors to this layer, making judgment calls the system routes to them (rather than reviewing outputs after the fact).

Cognition sharing in multi-agent systems can create new risks, including unintended delegations, malicious prompt injections or memory poisoning, or over-privileged agents with access to permissions and data far beyond what their tasks require. Environment-specific controls are therefore needed to protect against unintended actions or consequences.

“Agents have human-like attributes but operate at machine speed and scale,” says Pandey. “Everything we built for twenty years—access control, identity, compliance—was built for humans or machines, not both.”

Continuous Agent Semantic Authorization (CASA)—an open-source reference implementation developed by Outshift—is a GAT that works to ensure agent actions remain securely aligned with the user’s original goal through a process of continuous authorization. It does this by reading what the agent is trying to accomplish then checking each tool request against that task.

In the case of a healthcare system, for example, an agent told to summarize a patient record may start by querying a whole database. This could lead to CASA denying the call, because the request no longer matches the task it was authorized for.

“Today’s controls are scoped to a role or a session not to the task so an agent granted a tool can use it for anything,” explains Pandey. “Roughly 90% of the time, an agent has no way to confirm it is even cleared for the job it was handed.”

Experimentation for cross-domain innovation

When horizontally scaling intelligence in the enterprise, businesses should begin by experimenting with one cross-functional workflow that spans three or four teams and currently needs a human authorizing the handoffs, Pandey advises.

“Stand it up as a small multi-agent system on open, interoperable infrastructure, with a measurable baseline,” he says. “Keep building bigger models, add the horizontal axis on top of them, and change what you measure. Track where one agent’s insight made another agent better—that is the signal the horizontal axis is working.”

By starting to experiment now with intent, context, and reasoning layers, organizations can get ahead of the curve. “The problems are open, and the infrastructure is still being written,” says Pandey. “This is the moment to build it.”

For more information on the Internet of Cognition, visit Outshift.com.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Read more

Drug discovery is a high-cost, high-risk endeavor that is under growing pressure from a market increasingly defined by first-mover advantage.

Since the 1950s, the cost of developing new pharmaceuticals has roughly doubled every nine years—a phenomenon known as Eroom’s Law. Today, bringing a new drug to market takes an average of 10-15 years and costs anywhere from $1 billion to $2.5 billion, with failure rates upward of 90%.

AI has become the pharmaceutical industry’s biggest bet on bringing success rates up and timelines down. The faster drug companies can identify, test, and optimize new chemical compounds, the lower the risk of costly failures later in development.

“The main cost in drug discovery is still the clinical phase, so trying to reduce risk and increase your success rates there is obviously hugely beneficial,” says Paul Belcher, director of protein research strategy at global life sciences company Cytiva. “AI is one approach that drug companies hope will not only save time and compress timelines, but enable better quality candidates to reach the clinic.”

Early use of AI in drug discovery shows potential, but also highlights the need for robust and authentic data, as well as integration in lab systems.

AI brings efficiency to the lab

One of the most promising early-stage applications of AI in drug discovery is in hit identification. This involves screening libraries of molecular entities against a disease-related target, such as a protein, to find molecules that bind to it. A successful hit gives researchers a starting point for further testing and refinement, with the aim of eventually developing a viable drug.

Belcher has seen a shift from empirical screening to predictive design: Instead of physically screening libraries, drug companies are now using AI to design drug candidates from scratch and predict how they will interact with disease targets before committing anything to research and development (R&D).

This means companies are no longer limited by how much they can physically screen to identify starting points. “AI does away with that,” says Belcher. “And it can help eliminate low-quality candidates before you have to physically test them, saving time and resources.”

What AI can’t do yet is reliably predict kinetics or developability of new compounds, says Belcher. This means every AI-generated candidate still needs to be validated in the lab.

Traditional screening workflows were built to identify hits at scale, not to profile large numbers of complex candidates in detail. This is placing more pressure on lab teams, who now have to test, characterize, and purify a growing volume of more diverse, AI-generated compounds.

“The current techniques used in hit identification can screen hundreds of thousands, sometimes millions of compounds, using binary or threshold-based techniques producing low-fidelity data—yes-or-no responses,” Belcher explains. “AI can increase the number of hits you get and potentially give you better quality hits as well. That increases demand for higher-throughput, information-rich technologies to then validate and characterize those hits.”

Models need complete, quality data

As AI has accelerated demand for data-rich lab systems, it has also highlighted a fundamental need for better, more complete data.

Many earlier AI models were trained on publicly available datasets and are now hitting what Belcher calls a data wall. Because models have access to the same data, they all reach similar conclusions, with diminishing returns over time. Additionally, the datasets weren’t built with AI in mind, meaning they lack the structure, labeling, and diversity needed to keep models accurate and free of bias.

Publication bias reinforces the problem. “Most publicly available datasets and scientific publications focus exclusively on positive results,” says Belcher. “No one wants to share their failures. This bias is almost like having one hand tied behind your back. AI models can identify patterns associated with success, but they lack the comprehensive understanding of failures that would make predictions more reliable.”

The data Belcher believes would markedly improve models—the failed experiments, the compounds that don’t bind—remains frustratingly difficult to come by. “We often joke that there should be a journal of negative data,” he says. “It’s often buried in lab notebooks, and it’s never used to inform or guide future research.”

This lack of negative data creates a fundamental problem: Without access to a broad range of data, models can’t be adequately trained to avoid bias. “In all machine learning applications, the model’s performance relies heavily on the quality and scope of the training data,” notes Belcher.

Fabrication has also become much easier with AI, compounding concerns around data integrity. Take Western blots, for example. These are part of a standard technique for identifying proteins in blood or tissue samples, and they are among the most common targets for manipulation in biomedical research. Belcher cites research by Dutch microbiologist Elisabeth Bik, who found that almost 4% of biomedical papers contained duplicated or manipulated images. This was back in 2016, before generative AI made fabrication trivial.

“Manipulated or faked data has always been a problem in science, but in the AI world, especially when used to train models, it could have potentially disastrous consequences,” says Belcher. “There needs to be more tools to verify that data is not manipulated.”

Some vendors are starting to tackle this challenge. Belcher points to solutions like Cytiva’s Image Integrity Checker, for instance, which uses secure hash algorithms—the same technology used in blockchain—to detect whether scientific images have been tampered with. “We’re starting to see a lot of interest from publishing houses that want to adopt this as standard because it’s a quick way to ensure that what gets published in the literature is genuine,” he adds.

Autonomous labs could accelerate breakthroughs

Belcher describes the future state of drug discovery as fully autonomous labs that run with minimal human intervention. Foundational to this vision is consistency in data and infrastructure.

These AI-driven dark labs, or labs-in-the-loop, operate around the clock. They cycle through prediction, testing, and optimization, and then feed results back into AI models to guide the next round of experiments. This can improve the success rates of drug candidates entering clinical trials, says Belcher. Better starting points, combined with more rounds of optimization, should result in better candidates with fewer liabilities reaching the clinic.

But automating a lab depends heavily on integration. That means interoperable systems, highly structured and comprehensive datasets, and information flowing easily in and out. Most labs aren’t there yet. “Today, a lot of the instruments in labs are standalone,” Belcher notes. “You can have the best technology in the world, but if it’s a closed ecosystem—if the user can’t get the data out—it doesn’t do any good.”

An integrated infrastructure can enable labs to generate FAIR (findable, accessible, interoperable, and reusable) data at scale. This would not only inform individual lab reports, but could also train subsequent generations of AI models, effectively closing the loop between the computational, AI-driven dry lab and the physical wet lab.

“Our goal is to help scientists and researchers accelerate their breakthroughs and make that future state of autonomous labs a real possibility,” says Belcher. “We want to help them generate reliable data, simplify workflows in discovery, and hopefully enable what they’re working on to become tomorrow’s life-changing therapies, faster and with greater confidence.”

On costs and what comes next

AI-driven drug discovery is still in its early days. Notably, no drug discovered primarily through AI-driven design has yet received full FDA approval—although Belcher expects that to change in the next two to three years.

How big of an impact could AI eventually have on drug discovery? “The holy grail would be full in silico prediction of efficacy and toxicity, eliminating the need for the vast majority of physical wet lab work,” says Belcher. But there are many barriers to this beyond the maturity of the models, including regulatory hurdles and cost challenges.

A Stanford study found that the cost of training frontier AI models has more than doubled every year since 2016, adding more financial pressure to a sector already defined by exceptionally high R&D spend.

Belcher acknowledges the tension, but remains optimistic about what’s ahead. “I think we’ll get to a point where there’s a balance between AI and wet work, from a cost perspective and a risk perspective,” he says. “As long as the cost of compute doesn’t ever outweigh the cost of clinical development, I think AI is going to be an advantage.”

Learn more about how Cytiva is using faster discovery to reshape protein purification workflows.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Read more
1 … 80 81 82 83 84 … 3,356