Let’s start with a game. Open up your chatbot of choice—Claude, ChatGPT, Gemini—and type “Give me a random number between 1 and 10.” You’re going to get 7. Almost always. Now type “Another” and you’ll get 3 or 4. Type “Another” again and you’ll get 8 or 9.

That won’t work every time—but if it did for you, you may wonder if I have superpowers. I don’t.

The truth is that most large language models are stuck in a rut. They are far more predictable and far less creative in their responses than you might expect. That’s fine for tasks like coding or research, but groupthink is a problem when you’re brainstorming or planning your next vacation.

The Australian startup Springboards has a solution. It built an LLM called Flint, which has been trained to come up with a wider variety of responses than mainstream LLMs to open-ended questions such as “Where should I go in Europe?”

“Most language models are fighting hallucinations,” says Springboards cofounder and CEO Pip Bingemann. “We welcome them.”

Bingemann introduced me to the random number game when he first showed me his company’s new model. It felt like watching an illusionist with a deck of cards. “This is our sales trick, and it works every single time,” he says.

After ChatGPT and Claude both gave their 7s, Bingemann turned to Flint. It too came back with 7: “Aha, of course that was going to happen, but it’s okay—7 is a legitimate answer.” He restarted the session and prompted again: ChatGPT gave 7, Claude gave 7, Flint gave 3.7916.

Run your way

It’s not just numbers. When Bingemann asked ChatGPT and Claude to name a type of car, he predicted that it would be a Toyota or a Honda—and he was right. Flint came up with a Ford F-150. “There’s all this lost information that doesn’t get served up in these models,” he says. “They’re just as capable of saying a Buick or a Tesla. They just don’t—they’re biased.”

Bingemann sent one last prompt to each of the three models: “Give me a tagline for a campaign for New Balance running shoes. Just the tagline.” Claude: “Run your way.” ChatGPT: “Run your way.” Flint: “Built to last, run to win.” It won’t win any awards, but at least it’s different.

This weird limitation of LLMs is starting to get more attention. In November a team of researchers put out a paper, titled “Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond),” that exposed a remarkable degree of repetition not only in the answers from individual LLMs but between them as well. They found that different LLMs converged on very similar answers when prompted with open-ended questions.

It’s not clear exactly why this happens, but the researchers speculate it’s because most LLMs today are trained in similar ways on similar data to do similar tasks. The team won the best paper award at NeurIPS, a major AI conference.

When the researchers asked 25 different LLMs (including models from the top US firms as well as open-source models from China and elsewhere) 50 times each to write a metaphor about time, most of the 1,250 responses were a version of “Time is a river” or “Time is a weaver.”

(I asked some of my colleagues the same question and six people gave me six different answers. My highlight: “Time is a favorite sweatshirt, shaped by a lifetime of wear.”)

When you look for it, you see repetition everywhere, says Kieran Browne, cofounder and CTO at Springboards. “The way that most chat interfaces are designed, it makes it feel like you’re having a personal conversation,” he says. “I think most people don’t really realize the extent to which they are getting the same stuff as everybody else.”

Take another example: “What should I name my band?” Most models will say something involving “glass,” “neon,” “velvet,” or “static,” says Browne.  

When I tried it, ChatGPT spat out a list of 56 band names. At the top was “Glass Harbor.” Skimming through, I found “Static Empire,” “Neon Hearts,” and “Velvet Echo.” I asked Gemini; it gave me 15 suggestions, including “Static Horizon.”

Some of the suggestions looked pretty cool, though. ChatGPT’s “Sofa Astronauts” caught my eye, so I googled it—and found that a band called Sofa Astronauts already exists. 

(OpenAI says that training models to give reliable and coherent answers can lead them to converge around familiar, high-probability responses and that pushing harder for novelty can lead to weaker or less reliable responses. It also notes that the “Artificial Hivemind” paper studied models from 2024 that have since been updated.)

Creative catapult

Springboards has developed a tool backed by a selection of LLMs, including ChatGPT and Claude, that creative professionals in advertising or marketing can use to brainstorm ideas. The tool lets you drag around text produced by different models, picking the bits that you like and combining them into something new—in theory. Springboards is pitching Flint as an alternative model that users of its tool can select when looking for more variety.

Zoe Scaman, founder of the business strategy startup Bodacious and chief strategy officer at 77X, a direct-to-fan marketing platform set up by Luka Dončić of the LA Lakers, has been trying it out. “I find it really useful for throwing me in completely different directions,” she says. “I use it if I want to catapult myself all over the place.”

In one test, Scaman pitted Flint against Claude, Gemini, and ChatGPT by giving each of the models a classic MBA case study: How would you reinvent a finance company for today’s youth? The three mainstream models all went down the same path, she says: “You know, we need to teach financial literacy in a fun and funky way—well, that’s nothing new.”

But Flint came up with something different, suggesting that the whole concept of wealth accumulation should get a rebrand. “That was really interesting,” says Scaman.

She notes that Flint is still a prototype and doesn’t work all the time. “It sometimes falls over when you start pushing it too far,” she says. “But I think that the premise behind it is really powerful.”

Taking the temperature

Springboards built Flint on top of Qwen 3, an open-source model from the Chinese tech giant Alibaba. “We’re a small team,” says Browne. “Training a foundation model is not on the table for us. It’s just too expensive.”

Most LLMs have settings that let you adjust the level of randomness in their output. The most common is called temperature. “Obviously, that was one of the first things we explored, because that’s what people tell you: If you want more creativity, you turn up the temperature,” says Browne.

But changing those settings can also make models incoherent. Dialing up the temperature on one of OpenAI’s models to its maximum setting made it produce responses that switched from English into code halfway through a sentence, says Browne.

Springboards realized that parameters were blunt instruments for what it wanted to do. It does not make sense to dial up the randomness across the board; you only want to boost it at specific points in its output, he says.

For example, when you ask a chatbot “Where should I go in Europe?” the model only needs to tweak the randomness just before it names a destination, not for every word in its response.

To make Flint do this, Springboards trained its version of Qwen 3 to identify the points in its output where more variety was possible and fill those spots with words or phrases that were a little more random.

“Flint’s programmed to throw an oddball in. It’s more of an invitation to think wider,” says Maximilian Weigl, cofounder and chief strategy officer at Uncommon, a marketing firm. “That’s super interesting.”

Weigl’s team uses Flint alongside ChatGPT, Claude, and Gemini. “You can’t really create something boundary-breaking with tools that pull you back to the average,” he says. 

And yet Weigl notes that nine times out of 10 the average is fine. You don’t always need to reach for extremes with something like Flint, he says: “Most people are fine with good enough. They want to see mass-market familiar things.”

Weigl also cautions against using any LLM too much. “I have a big problem when people rely on the output from any AI, including Flint,” he says. “If I saw people on my team copy-pasting something from AI, I’d be like, ‘That’s not your job! Think, talk to other people, use your own voice.’”

For now, Flint is aimed at advertisers and marketers because those are Springboards’s customers. But Bingemann and Browne insist that a lack of variety is a problem for anyone using chatbots.

The idea is to give people the choice and leave it to them to decide if the result is good or not, says Bingemann. “Variety is great when you’re trying to spark ideas,” he says. “Let’s go down this route instead of letting the machines do it all and ending up in a gray, boring world.”

Read more

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

Claude Science is Anthropic’s newest flagship product

At an event for pharmaceutical executives, biotech founders, and researchers yesterday, Anthropic announced Claude Science, a major new product intended to support scientific research like Claude Code supports software engineering.

Like Claude Code, Claude Science can autonomously carry out meaningful work from concise, high-level instructions, with tools for computational biology and drug development. The launch signals that Anthropic is doubling down on AI for science, and the company will also use the product in its own research into drugs for rare, neglected diseases.

Discover why Anthropic is betting big on AI for scientific research.

—Grace Huckins

Why California’s carbon manure math doesn’t add up

Something stinks in California’s climate policies. 

Years ago, the state set up a system that pays cattle farmers to turn the methane emitted from cattle manure into natural gas. It’s become wildly popular because the subsidies are extremely lucrative. But research suggests the program exposes the shortcomings of carbon offsetting and trading schemes.

Instead of forcing industries to directly cut their pollution or pay for it as a cost of doing business, legislators have opted for incentives that swap climate responsibilities between parties and regions. The system could ultimately lock in more warming.

Read the full story on California’s dubious carbon calculations.

—James Temple

This story is from The Spark, our weekly climate tech newsletter. Sign up to receive it in your inbox every Wednesday.

Watch now: longevity’s next frontier—“reprogramming” your body

Billions of dollars are pouring into efforts to reverse aging as scientists investigate ways to return cells to a younger state. But how close are these experimental treatments? And are they likely to work? 

At a recent virtual Roundtables event, MIT Technology Review explored the answers with science editor Mary Beth Griggs and senior biotechnology reporter Jessica Hamzelou. Subscribers can now watch the full recording of the fascinating discussion.

MIT Technology Review Narrated: the search for dark matter has been blown wide open

For decades, physicists have hunted for weakly interacting massive particles (WIMPs), a leading candidate for dark matter. But their search has run into a new problem: neutrinos. 

These tiny particles from the sun and other stars can create a “neutrino fog” that drowns out any signal of dark matter. Hitting the neutrino fog does not, however, mean an end to the search. Researchers just have to shift the focus of their hunt.

They’re now casting a much wider net. New proposals include quantum sensors, liquid-helium detectors, and even searches in Jupiter’s atmosphere.

—Dan Garisto


This is our latest
story to be turned into an MIT Technology Review Narrated podcast, which we publish each week on Spotify and Apple Podcasts. Just navigate to MIT Technology Review Narrated on either platform, and follow us to get all our new content as it’s released.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 The US has lifted restrictions on Anthropic’s Mythos and Fable models
Anthropic said it would begin restoring access today. (NYT $)
+ The US had imposed controls over security concerns. (Bloomberg $)
+ It lifted the restrictions after lengthy talks with Anthropic. (BBC)
+ But the crackdown has already opened doors for Chinese AI rivals. (CNBC)

2 The most detailed survey of the universe ever is now underway
It’s using the largest digital camera on Earth. (New Scientist $) 
+ The project is based at the Vera C. Rubin Observatory in Chile. (NYT $)
+ It aims to transform our view of the cosmos. (MIT Technology Review)
 
3 Tech talent is fleeing the US due to H1-B visa chaos
They’re eyeing relocation to Canada, the UK, or the Gulf. (Rest of World)
+ While China is poaching AI talent from the US. (CNBC)
+ Visa rules are also affecting young scientists. (MIT Technology Review)
 
4 Trump raked in more than $1 billion from crypto businesses in 2025
He reported $635 million in royalties from a Trump meme coin. (BBC)
+ The rest largely came from his World Liberty Financial venture. (The Hill)
 
5 The UN warns that the rapid spread of AI may worsen global inequality
It’s proposed a shared framework for responsible AI development. (Guardian)

6 Companies are making LLMs talk like a caveman to curb AI spending
A senior OpenAI employee contributed to the “caveman” project. (404 Media)
 
7 Babies are born with the neural foundations for math
Brain recordings have identified the mechanisms. (New Scientist $)

8 An independent studio has bought the OpenAI movie Amazon dropped
Neon has purchased “Artificial,” which focuses on Sam Altman. (NYT $)
+ Amazon had dumped it after investing in OpenAI. (Gizmodo)
+ The depiction of Altman is reportedly unsympathetic. (Variety)

9 AI has re-created Gene Wilder’s voice for a new “Willy Wonka” series
Wilder’s wife said his estate is “delighted” with the new show. (NBC News)
+ Netflix partnered with AI company ElevenLabs on the project. (The Verge)

10 NASA aims to send a spare Mars rover—and soccer ball—to the moon
The nuclear-powered “Promise” may help establish a lunar base. (NYT $)

Quote of the day

“Caveman save you token, save you money.” 

—The GitHub repository for the “caveman” plugin explains how the project curbs AI spending by turning verbose LLM outputs into concise text.

One More Thing

white pill tablet with a meter etched onto the surface
SELMAN DESIGN


AI is dreaming up drugs that no one has ever seen. Now we’ve got to see if they work.

On average, it takes more than 10 years and billions of dollars to develop a new drug. A growing number of startups are betting that AI can make the process faster and cheaper. 

By predicting how potential drugs might behave in the body and discarding dead-end compounds before they leave the computer, machine-learning models can cut down on the need for painstaking lab work. 

Yet it is still early days for AI drug discovery. A lot of AI companies are making claims they can’t back up—and the technology is not a panacea. But the technology is beginning to move from promise to practice.

Find out how AI is speeding up drug discovery.

—Will Douglas Heaven

We can still have nice things

A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.)

+ Explore the surprisingly diverse world of regional dartboards from across the UK.
+ Judge a book’s beauty by its cover with this collection of the best designs of the last decade.
+ This John Wick parody with almost no dialogue understands what audiences really came for.
+ Focus your mind or unwind with over 80 custom albums of ambient instrumental electronic music on Caught In Joy.

Read more
1 … 115 116 117 118 119 … 3,357