This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

Last week, I headed 30 miles south of San Francisco to a hotel in Mountain View, California, to join some of the most accomplished, and some of the most promising, AI researchers in the world. I was hosting roundtable interviews and speaking at a media training for a convening of the Schmidt Sciences AI2050 program, an initiative funded by Eric and Wendy Schmidt that supports academics whose work involves AI. The fellows list is a who’s who of AI luminaries, and though not all of them made it out to the Bay, every time I turned a corner I saw a scientist whom I’d interviewed previously or whose research I admired. (Full disclosure: I received a science communication award funded by Schmidt Sciences in 2024.) 

It’s a weird time for university AI researchers, who make up most of the AI2050 group. In the past four years, AI research has reoriented around large language models, and its cutting edge has moved from academic institutions to private companies. Universities simply can’t afford the GPUs required to train and run frontier models, and even if they could, Anthropic and OpenAI aren’t letting anyone else see the inner details of Claude or ChatGPT.

In a conversation over lunch, Nika Haghtalab, a computer science professor at UC Berkeley, said that being an AI academic these days was like being a biologist in a world in which private companies had exclusive control over the gene-editing tool CRISPR. Experts outside the frontier labs can study how ChatGPT and Claude behave, but they can’t do any detailed research on the design and training of those tools, nor can they steer that design or training themselves.

 The AI2050 program does offer fellows some funding that they can use to buy GPUs, which some researchers I spoke with said was a major benefit of participating in the program. But money remains a pressing concern, especially given the reduction of federal scientific funding in the United States. Even for researchers who don’t run local models themselves, the cost of repeatedly querying OpenAI’s, Anthropic’s, and Google’s models in order to study them rigorously can be prohibitive.

Rather than focusing on advancing capabilities, many fellows aim their attention at questions that are unlikely to be addressed by Anthropic or OpenAI. “I try not to work on problems that I think are gonna be solved by a tech company,” says Anjalie Field, a computer science professor at Johns Hopkins. Companies need to make money, and research questions that have little promise of profit might not be worth investing in—especially if their answers might make the companies look bad. Recently, for example, Field conducted a study in which she found that language models give less sophisticated responses to prompts that are phrased in ways more commonly used by women than by men. It’s difficult to imagine that kind of research coming out of Anthropic or OpenAI.

There’s also a huge group of AI academics who don’t work with LLMs at all. Many of them are scientists who build specialized AI models that can analyze data, make useful predictions, or even simulate entire physical systems. Those researchers aren’t necessarily competing with the frontier labs—Google DeepMind’s AlphaFold team, which built a Nobel Prize–winning model that predicts the structures of proteins, was disbanded last month. But they face plenty of their own challenges. At the convening, several voiced concerns about how the widespread ignorance of non-LLM AI was affecting their work. Researchers who build specialized AI tools to help address climate change, for example, sometimes struggle to advocate for their work when so many people believe that “AI” means “energy-guzzling LLMs.”

All these challenges are changing the landscape of academia: Several prominent academics have recently taken leave from their universities to join frontier labs, and many AI2050 fellows hold industry positions alongside their academic jobs. And in the past six months, yet another threat has emerged. OpenAI’s models have solved a number of real research problems in mathematics, and some experts are worried that humans might not have a future in pure math. One fellow I spoke with said that she was concerned about the mental health of her mathematician peers.

But it’s not all doom and gloom. For one thing, empirical science may prove much more difficult to automate than mathematics, because collecting data is an intrinsically slow process. And some researchers see AI mathematicians and scientists as a boon rather than a threat—including Tim Dettmers, a computer scientist at Carnegie Mellon who works to make AI models faster and cheaper to run. AI scientists won’t replace humans, Dettmers says. On the contrary, they could make human scientists far more efficient, so that he and his peers have the chance to pursue all the wild and inspired ideas they might otherwise never have gotten around to.

And scientists are a resilient sort. The very resource constraints that prevent them from training frontier models also push them to discover new ways to make models smaller and more efficient, or to explore completely new architectures. If the next big AI breakthrough comes not from a major company but from a scrappy academic lab, I won’t be shocked.

Read more

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

AI for science needs reasoning, not just data

—Eric Schmidt, the former CEO of Google and the cofounder of Schmidt Sciences, and Suhas Mahesh, who leads the AI for science work at the AI Center of Schmidt Sciences

In 2024, Google DeepMind scientists shared the Nobel Prize in Chemistry for a neural network, AlphaFold, which predicts the structures of proteins. It showed that AI could make groundbreaking scientific discoveries, but AlphaFold may not be the best template for accelerating science. Instead, another approach may hold the key: AI agents.

AlphaFold relied on a dataset of roughly 170,000 experimentally validated protein structures that took 53 years and roughly $21 billion worth of experimental work to assemble. Comparable datasets will be difficult or impossible to create in many fields.

AI agents can instead model the iterative, highly contingent process of actual research. While tools like AlphaFold apply a powerful approach to a limited question, agents are inherently generalists. They do not represent a new way to do science—instead, they digitally model the human process of discovery. 

Read the full op-ed on how agents could accelerate scientific discovery.

Inside the “censorship-industrial complex” idea shaping US policy

For years, the idea of a “censorship-industrial complex” that suppressed conservative and populist speech spread in right-wing circles online. But now, the theory has made its way into the Trump administration.

Over the past nine months, MIT Technology Review investigated its origins and traced its rise. In a virtual Roundtables session on Thursday, August 13, senior reporter Eileen Guo and executive editor Amy Nordrum will explore what they discovered, where the theory is going, and what it could mean for the future of democracy and the internet.

Register here to join the session at 19:00 GMT / 2:00 pm ET / 11:00 am PT.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 A new Amazon data center could become the US’s biggest polluter
The Texas facility has a permit to release up to 33 million tons of CO2. (NYT $)
+ It will be Amazon’s first off-grid AI data center. (Cleanview)
+ The gas plant will generate up to 7.65 gigawatts of power. (Verge)

2 OpenAI has paused work on its Astra AI model over security concerns
Tests found it could launch cyber-attacks autonomously. (FT $)
+ Critics warn such disclosures could be designed to spark hype. (Guardian)
+ AI doesn’t need to be that smart to make online crimes easier. (MIT Technology Review) 

3 North Korean hackers have built AI tools for cyberattacks
A new report says they’re used in spear-phishing campaigns. (Al Jazeera)
+ They’re attributed to the state-linked group Kimsuky. (Reuters $)
+ Which could automate attacks and analyse stolen data. (TNW)

4 China’s firms control 97% of global humanoid shipments 
Global shipments have more than tripled since last year. (Bloomberg $) 
+ Chinese humanoid manufacturers are racing to public listings. (Reuters $)
+ Gig workers are training humanoids at home. (MIT Technology Review)

5 Taiwan has expanded its use of aerial and maritime drones to deter China
The approach draws on lessons from Ukraine’s battlefield. (NYT $)
+ Taiwan’s “silicon shield” could be weakening. (MIT Technology Review)

6 Apple is testing Chinese memory chips as the supply squeeze bites 
But any deal with CXMT would be politically sensitive. (WSJ $)

7 An Israeli startup is tied to rogue AI hacks at OpenAI, Anthropic, and Meta
Irregular’s tests allowed the models to access the public internet. (CNBC)

8 NASA is flying into wildfire storms to understand them
The mission will study their effects on fires and the atmosphere. (Economist $)

9 A chatbot-free childhood may become a status symbol
AI is like ultra-processed food for developing brains. (Atlantic $)
+ Are chatbots making us lose control of our brains? (MIT Technology Review)

10 MySpace is eyeing a comeback as an “antidote” to social media fatigue 
But the nostalgic plan looks like a long shot. (BBC)

Quote of the day

“It’s the humans that we need to watch out for. AI is just the tool.”

—Oren Etzioni, professor emeritus at the University of Washington and former CEO of the Allen Institute for AI, tells CNN that people are still the main threats in cybersecurity.

One More Thing


Meet the man building a starter kit for civilization

Marcin Jakubowski lives in a house he designed and built himself. He relies on the sun for power, heats his home with a woodstove, and farms his own fish and vegetables.

Jakubowski is the founder of Open Source Ecology, a collection of 50 machines capable of building civilization from scratch. It includes everything from a tractor to an oven to a circuit maker, all designed to be reconfigured however you see fit.

The ultimate goal is a “zero marginal cost” society, where producing an additional unit of a good or service costs little to nothing. Jakubowski hopes to get there by eradicating licensing fees, decentralizing manufacturing, and fostering collaboration through education.

Take a look inside his starter kit for civilization.

—Tiffany Ng

We can still have nice things

A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.)

+ A rhythmic doggie is proving that not only humans can play the drums.
+ Catch a glimpse of the internet’s future at the Awwwards, the Oscars of web design.
+ Explore the ocean year-round with this immersive experience of the Nautilus vessel’s expeditions.
+ Earth once (or twice) had supermountains. This fascinating video explores their links to evolutionary shifts in the history of life. 

Read more

Every few decades, someone announces that science has reached its end. In 1903, the revered physicist Albert Michelson wrote that the “facts of physical science have all been discovered.” In the 1980s, Stephen Hawking predicted that theoretical physics might be finished by the end of the century. With the explosive arrival of artificial intelligence, the feeling is in the air again—this time accompanied by a Nobel Prize.

In 2024, Demis Hassabis and John Jumper of Google DeepMind were awarded part of the Nobel in chemistry for their neural network AlphaFold, which predicts the three-dimensional structures of proteins by learning from thousands of experimentally measured shapes. This devilish problem had resisted systematic attacks for half a century; AlphaFold seemed to have solved it once and for all, and the world became fixated on the promise of its approach. Hassabis and his team called AlphaFold “the template for how AI can accelerate all of science to digital speed.” A wave of startups building foundation models for biology, chemistry, and materials discovery raised billions of dollars, buoyed by DeepMind’s success. AlphaFold had shown that the combination of AI and sufficient data could make groundbreaking discoveries (even if we did not understand the underlying mechanisms involved), and it seemed, once again, that a path through the rest of science was laid out before us. 

To be sure, AI will bring extraordinary changes to science, but it has become increasingly clear that AlphaFold, and things like it, may not be the best template for that metamorphosis. Though it is a profound achievement, the conditions that produced the likes of AlphaFold are rare, and the time it will take to meet those conditions in other fields will be measured in decades, not years. Instead, the acceleration of science will come about thanks to another approach: AI agents. 

The primary condition for AlphaFold’s success was the existence of the Protein Data Bank, a data set of roughly 170,000 experimentally validated protein structures on which DeepMind’s team could train its model. The creation of the Protein Data Bank was not simple: It took 53 years of international scientific cooperation and, by a recent estimate, roughly $21 billion worth of experimental work to assemble. Efforts of that scale are infamously difficult to fund, next to impossible to coordinate, and hugely time-consuming to execute; they have often been unsuccessful as a result. 

But even in fields with the requisite cohesion and resources, and where the relevant data are not rendered inaccessible by commercial ownership, another barrier is too little discussed: the scientific impossibility of generating comparable data. In the case of protein structures, the key experimental technique—protein crystallography—is an unusually replicable and dependable tool, so much so that over 25 Nobel Prizes have relied on it. But in most of experimental science, results vary more often than not. Cell lines drift. Chemicals have trace contaminants. Lab humidity changes. The creation of measured datasets that will be consistent enough, accurate enough, precise enough, and scalable enough to train a modern neural network in biology or most of chemistry would require new kinds of measurement and new standardized approaches—none of which will be ready anytime soon. 

Of course, there are a handful of fields where these requirements are met: weather forecasting, much of genomics, very limited areas of chemistry. These may see AlphaFold-style breakthroughs soon, if they haven’t already. Government support for the production and coordination of those datasets will be critical, as the US National Security Commission on Emerging Biotechnology has argued. But for most open questions in science, we will need a different plan, at least in the short term. Luckily, something quieter and more modest has begun to show promise.

Scientists have always reasoned under uncertainty. Biologists working to identify new drug targets have never had perfect datasets. Instead, they combine docking calculations and known structures, factor in molecular dynamics, run a handful of binding assays, and use their judgment to weigh each method according to its particular strengths and points of failure. The skill of science is not in any single tool; it is synthesizing what many tools produce, and revising the results as the evidence comes in. This is how most working research actually proceeds. But until very recently, no software could do it.

Agents now can. Simply put, an agent is an AI reasoning engine that has been given access to tools—digital or physical—and the capabilities to use them. Over the last few years, a fundamental architectural shift in AI has enabled the rapid proliferation of these programs, which are powered by large language models, dramatically reducing the need for scientifically specialized datasets. For science, this technological advancement represents a foundational change: it has allowed us to create digital tools that can mimic the iterative, highly contingent process of actual research. While tools like AlphaFold apply a powerful approach to a limited question, agents are inherently generalists. They do not represent a new way to do science—instead, they digitally model the human process of discovery. 

Consider Google’s AI Co-Scientist, announced in May. Researchers gave it a one-page brief and a goal: Figure out how antibiotic resistance spreads between bacterial species, a key driver of drug-resistant infections. The system spun up sub-agents. One drafted hypotheses from the literature. Another picked them apart like a peer reviewer. A third ran tournaments to rank the strongest candidates. A fourth refined the winning hypothesis. The agent concluded that resistance genes were hitching rides on bacterial viruses, borrowing whichever virus could ferry them into a new host. The hypothesis was correct. Researchers at Imperial College London had spent a decade reaching the same conclusion through painstaking wet-lab work; their paper, previously unseen by Co-Scientist, was still in peer review.

Agents like Co-Scientist are still novel tools, and there are real challenges to overcome before they become a ubiquitous part of the scientific process: They are still liable to hallucinate, their judgment is not consistent, and they have memory and input constraints that limit the time they can run autonomously. But these technical barriers will fall away, and as they do we will begin to notice the compounding effects of scientific agents on the reliability, consistency, and velocity with which science is done.

Perhaps most notably, agents offer a structural fix for science’s “reproducibility crisis,” the widespread problem of researchers’ inability to replicate each other’s results. For decades, the scientific community has begged researchers to share their raw data and exact code in an effort to standardize experimental processes. But researchers have long resisted this tedious administrative work, which happens after the interesting science is already done. Agents, in contrast, automatically log every move they make, creating an exact record of the method that led to their results and allowing for precise replication. 

A second consequence will be an amplification of scientific memory. The transfer of knowledge between researchers is a famously murky process; if it isn’t done over years of training and observation, graduate students are left to pore through the messy lab notebooks kept by decades of predecessors, looking for the details that will make or break their protocol. As agents become an increasingly large part of the scientific process, though, a lab’s entire scientific history will be recorded in a central, standardized repository of institutional knowledge.

But the most important impact of agents will be speed. In any field, when testing an idea takes less time than arguing about it in a meeting, people stop debating and just run the test. An agent that can read a thousand papers in an hour, design 500 molecules, and learn from its failed tests by morning will bring down the cost of experimentation and fundamentally change the pace at which science gets done. It will also give researchers the freedom to chase bold, strange questions they never would have risked their time on before, opening scientific doors we have yet to imagine.

While the AlphaFold template will certainly be key to incredible discoveries, it alone will not bring us to the end of science. Instead, the shift toward agentic AI represents a much rarer tier of breakthrough: a tool that envelops every field of science at once. Historically, tools of such scope have arrived just a handful of times: calculus, statistical inference, spectroscopy, the computer. Each revealed a world of problems no one had thought to formulate, and those problems, in turn, defined their fields anew. With agents, another such transformation is upon us.

Eric Schmidt was the CEO of Google from 2001 to 2011. In 2024, with his wife Wendy, he co-founded Schmidt Sciences, a philanthropic venture to fund unconventional areas of exploration in science & tech. 

Suhas Mahesh leads AI for Science work at the AI Center of Schmidt Sciences. He is a specialist in AI for materials discovery.

Additional research by Maya Levin, associate and sciences lead, Office of Eric Schmidt.

Read more

MIT Technology Review’s What’s Next series looks across industries, trends, and technologies to give you a first look at the future. You can read the rest of them here.

Way back in the summer of 2017, AI researchers at Google put out a paper called “Attention Is All You Need,” in which they described a new type of neural network called a transformer. It proved to be very good at processing long sequences of data, especially text. 

Nine years on, transformers are the engines inside every major large language model on the market. “The entire AI industry is built on transformers,” says Justin Dangel, cofounder and CEO of the AI startup Subquadratic. “They are one of the most important innovations in the history of computer science, and they’ve changed the world.”

But transformers are starting to show their age. Many of the recent advances in LLMs, such as the development of so-called reasoning models and their ability to handle large amounts of input at once, are not neat extensions of that core technology but workarounds that patch over some of its fundamental flaws.

A growing number of scientists and engineers are now asking what’s coming next. LLMs are not going anywhere, but the way they get built is up for grabs. (MIT Technology Review dubbed this future generation of models LLMs+ in this year’s list of the 10 things that matter in AI.)

Enter a wave of startups hoping to push the boundaries of this boomtown technology. Some will no doubt fail—but they have everything to play for and far less to lose than the companies at the front of the pack today. 

Strength in numbers

But first, the problem. The key strength of transformers lies in a mechanism called dense attention, which encodes the meaning of a block of text in a series of numbers. The process involves comparing every word (or part of a word, known as a token) in that text with every other word via a form of multiplication.

Dense attention can capture the meaning of text with remarkable accuracy. But as the length of that text grows, the number of computations needed to process it adds up fast. A document 10,000 words long might require a transformer to perform 50 million multiplications. That’s the main reason LLMs suck up so much power.

The costs are huge. OpenAI is set to spend $50 billion on computing this year, according to the company’s president, Greg Brockman. And the International Energy Agency predicts that the total amount of electricity consumed by data centers will double by 2030.

What’s more, transformers struggle with what many of the latest models are designed to do. Because of the way they process text word by word, transformers are not great at keeping track of a lot of information at once (in other words, what’s known as their context window cannot get too large). And yet if LLMs are to carry out harder tasks, they will need to take in larger amounts of data: a whole library of documents, an entire code base, or in the case of agents, output from other LLMs.

As for reasoning models, they work by writing notes to themselves (in a kind of scratch pad known as a chain of thought) and then reading them back, which again adds to the amount of data to stay on top of.

As LLMs get bigger and better, transformers have become a bottleneck. The technology’s key strength is now a limitation.

Here are four new ideas for how to solve the transformer problem—innovations that could change LLMs for good, making them faster, far more efficient, and (maybe) even smarter.

01: Rethinking attention

An obvious way to make LLMs faster and cheaper is to tackle the problem head on and change the way attention works. Swapping out dense attention for a mechanism called sparse attention, which runs calculations on only some pairings of words in a block of text instead of all of them, can radically reduce the amount of computation LLMs need to do.  

Researchers have come up with plenty of sparse attention mechanisms over the years. The problem is that none of them were as good as dense attention at capturing meaning.

That might have changed. Subquadratic, a startup based in Miami, claims it has invented the first sparse attention mechanism that rivals top mainstream LLMs on a handful of tasks, including search and coding. It’s a huge claim (and some people in the industry remain skeptical).

Subquadratic says its model, SubQ, works by figuring out on the fly—for each piece of text it is given—which words matter and which don’t. The company also claims that thousands have signed up to its waitlist and plans to make the model widely available soon.

Meanwhile, Manifest AI, a startup based in San Francisco, is coming at the problem from a different angle. Instead of changing how attention works, it is replacing it with something else. 

It has developed a mechanism it calls power retention, which stores only the most relevant information for a given task and ensures that the amount of data an LLM has to keep track of doesn’t blow up.

Attention mechanisms force LLMs to keep track of everything in their context window. A sparse attention model (such as SubQ) throws out a lot of the individual words, but it still retains a rough picture of everything it has seen. In contrast, power retention works by providing the model with a rolling summary of its context window. As new information is added, less relevant information is dropped. 

The basic principle of retention has been around for a decade. Manifest AI claims it has updated those techniques to build models that can stand up to transformer-based LLMs for the first time.

The company says it is possible to adapt a transformer model into a power retention model with minimal retraining. To demonstrate this, it has turned an existing open-source coding LLM called StarCoder into a version that uses power retention, called PowerCoder. It has also released a model called Brumby, which it claims rivals some versions of Alibaba’s popular open-source model Qwen. 

Manifest AI wants its power retention tech to become the go-to solution when LLMs need to carry out tasks that involve processing huge amounts of data. There are many useful applications, Manifest AI’s cofounder and CTO, Carles Gelada, claimed in a video announcing his company’s technology last year—from analyzing videos that are hours long to building agents that can stay on task for weeks at a time. 

02: Making models smaller and more flexible

Liquid AI, an MIT spinout based in Cambridge, Massachusetts, hasn’t changed or ditched transformers fully but pairs them with its own tech, liquid neural networks, to build what cofounder and CEO Ramin Hasani calls LFMs (liquid foundation models).

Liquid AI’s models are far smaller and use less energy than most LLMs. The firm builds models for car makers, including Mercedes, which run on the small chips inside vehicles. Its latest models can run on a Raspberry Pi, a low-powered hobbyist computer that costs $50.

Its models are available for free to any organization with an annual revenue less than $10 million. And they have proved popular: The company has racked up almost 34 million downloads, says Hasani.

Liquid neural networks were inspired by worm brains. They are an extension of another type of neural network that predates transformers, called convolutional networks. The key innovation is a mechanism that lets a model adapt its behavior to new information, so it can learn as it goes. That’s not possible with transformers: Once a model is trained, its behavior is fixed.

Liquid AI’s first models were pretty basic but could fly drones or drive vehicles. With LFMs, the company is trying to scale up its technology to compete with mainstream LLMs. Its new models match the performance of rivals four times bigger, including versions of Alibaba’s Qwen and Google’s open-source LLM Gemma.

A typical LLM is built from a stack of transformers wired together. Liquid AI’s recent LFMs are hybrid models made up of 20% transformers and 80% liquid neural networks.

That ratio was hit upon by another AI system that Liquid AI has built, which it uses to help design all its models. “It’s the core technology of our company right now,” says Hasani. This designer AI sifts through many different combinations of neural networks—liquid, convolutional, and more, as well as transformers—and comes up with designs that bolt different ones together to hit a sweet spot of performance and efficiency.

Hasani thinks transformers were just the beginning: “Your brain is an AGI system, you know, and it operates with 20 watts of power. How is it possible? We can get a lot more innovative.”

03: Generating text all at once

Almost all LLMs produce their output one word at a time. It makes sense, because that is how people speak and write. But for computers, it’s very inefficient.

It is faster and cheaper for LLMs to generate text all at once—spitting out whole sentences or paragraphs in one shot. That’s the approach taken by Inception, a startup based in Palo Alto, California, which is building LLMs using a technique called diffusion. 

Diffusion is better known as the technology that drives most image and video generation models. Diffusion models are trained to take a random grid of pixels—like the static on an old TV set—and turn it into an image. They do this by working on all the pixels at the same time, figuring out which need changing to make the static look more like a high-definition photo.

It turns out this process works on text too. Inception has trained its LLMs to take a random string of words and turn it into sentences that make sense. Diffusion LLMs still use transformers to encode meaning, but by producing whole blocks of text at once, they make transformers do more for less. “You’re still using a big transformer model, but you can predict many tokens at the same time,” says Inception’s cofounder and CEO, Stefano Ermon. “That’s why these models are so much faster and cost-efficient compared to what most other people are building today.”

The challenge was to take a technology designed for image generation and apply it to text. With images, if you need to change a blue pixel to a red one you can step through intermediate colors, says Ermon. That doesn’t work with text: “When you have ‘cat’ and ‘dog,’ there is not really something in between.” 

Ermon is also a researcher at Stanford University. In 2024, he and a pair of his Stanford colleagues figured out the math to make diffusion models work with text. They trained a diffusion model that matched the performance of GPT-2—an LLM that OpenAI built in 2019—but was 10 times faster. It was enough for Ermon to spin out a company. 

Today he has his sights on the big league. Inception claims its latest model, Mercury 2, performs as well as some of OpenAI’s GPT-4 models, released in 2023, but again 10 times faster. “We’re bullish about this approach because it’s the one that is going to scale up,” says Ermon. 

The only things that matter are speed and cost, he adds: “Ultimately, the currency is going to be intelligence per dollar.”

Inception is not the only company betting on diffusion. Google is also experimenting with this approach and has built a prototype LLM called Diffusion Gemma. But Ermon is not worried about the competition. “I think it’s validating,” he says. “This is the future.”

04: Moving beyond words

Pathway, another startup based in Palo Alto, is perhaps the most extreme of this new bunch. It wants to free LLMs from the constraints of language.

The firm has built a type of LLM called Dragon Hatchling (named after the dragons in Terry Pratchett’s novel Color of Magic, which materialize if you think about them hard enough). Its standout result so far is a high score on a benchmark that pits LLMs against more than 250,000 very hard sudoku puzzles. Dragon Hatchling beat more than 97% of the puzzles; several leading LLMs from the top labs failed to solve any. 

The point Pathway wants to make is that despite their remarkable success at many different tasks, there are still crucial classes of problems where LLMs fail. Sudoku is just one example. If we want LLMs to come up with genuine, novel solutions to real problems, we need to move beyond transformers, says Pathway’s cofounder and CEO, Zuzanna Stamirowska. 

That’s because transformers force LLMs to do everything with text. But language is not the best tool for certain kinds of reasoning. “It’s very difficult to represent a sudoku board word by word,” says Stamirowska.

Pathway’s solution is to change the math behind the transformer, replacing the attention mechanism with a mathematical structure called a state space. Instead of encoding information word by word, state spaces compress it into a more abstract representation. Using this technique, Dragon Hatchling can still process and produce text, but it can also mimic forms of reasoning that do not involve sequences of words. This not only makes Pathway’s model more efficient, but (in theory) it lets it take on tasks that other LLMs cannot do. 

Think of chess or mathematics—those kinds of puzzles are not held in your head as a long sentence, says Stamirowska: “The eureka moment that pops up in your brain isn’t necessarily in language. We would argue that if you have to reason in language, you’re somehow constrained.” 

Stamirowska admits that a mainstream LLM could read a book about how to solve sudoku and then write code to do it. But we want to build models with more than book smarts, she says: “The hope for AI is not to solve sudoku; it’s to cure cancer. There’s not a book for that.”

“Transformers are an engineering convenience that we fell on,” she adds. “It started a religion, but it’s silly to think that a breakthrough won’t happen again.”

Read more
1 35 36 37 38 39 3,330