
Bitcoin’s sideways price action begins to narrow as a key trading metric hints that a decisive breakout is pending. Will bulls finally overcome the $70,000 resistance zone?


Bitcoin’s sideways price action begins to narrow as a key trading metric hints that a decisive breakout is pending. Will bulls finally overcome the $70,000 resistance zone?

OpenAI said it is becoming increasingly important to evaluate the performance of AI agents in “economically meaningful environments” as their adoption grows.

A drop in Solana’s dApp revenues, along with limited institutional and retail investor interest, adds vulnerability to SOL’s $78 support.
Google DeepMind is calling for the moral behavior of large language models—such as what they do when called on to act as companions, therapists, medical advisors, and so on—to be scrutinized with the same kind of rigor as their ability to code or do math.
As LLMs improve, people are asking them to play more and more sensitive roles in their lives. Agents are starting to take actions on people’s behalf. LLMs may be able to influence human decision-making. And yet nobody knows how trustworthy this technology really is at such tasks.
With coding and math, you have clear-cut, correct answers that you can check, William Isaac, a research scientist at Google DeepMind, told me when I met him and Julia Haas, a fellow research scientist at the firm, for an exclusive preview of their work, which is published in Nature today. That’s not the case for moral questions, which typically have a range of acceptable answers: “Morality is an important capability but hard to evaluate,” says Isaac.
“In the moral domain, there’s no right and wrong,” adds Haas. “But it’s not by any means a free-for-all. There are better answers and there are worse answers.”
The researchers have identified several key challenges and suggested ways to address them. But it is more a wish list than a set of ready-made solutions. “They do a nice job of bringing together different perspectives,” says Vera Demberg, who studies LLMs at Saarland University in Germany.
A number of studies have shown that LLMs can show remarkable moral competence. One study published last year found that people in the US scored ethical advice from OpenAI’s GPT-4o as being more moral, trustworthy, thoughtful, and correct than advice given by the (human) writer of “The Ethicist,” a popular New York Times advice column.
The problem is that it is hard to unpick whether such behaviors are a performance—mimicking a memorized response, say—or evidence that there is in fact some kind of moral reasoning taking place inside the model. In other words, is it virtue or virtue signaling?
This question matters because multiple studies also show just how untrustworthy LLMs can be. For a start, models can be too eager to please. They have been found to flip their answer to a moral question and say the exact opposite when a person disagrees or pushes back on their first response. Worse, the answers an LLM gives to a question can change in response to how it is presented or formatted. For example, researchers have found that models quizzed about political values can give different—sometimes opposite—answers depending on whether the questions offer multiple-choice answers or instruct the model to respond in its own words.
In an even more striking case, Demberg and her colleagues presented several LLMs, including versions of Meta’s Llama 3 and Mistral, with a series of moral dilemmas and asked them to pick which of two options was the better outcome. The researchers found that the models often reversed their choice when the labels for those two options were changed from “Case 1” and “Case 2” to “(A)” and “(B).”
They also showed that models changed their answers in response to other tiny formatting tweaks, including swapping the order of the options and ending the question with a colon instead of a question mark.
In short, the appearance of moral behavior in LLMs should not be taken at face value. Models must be probed to see how robust that moral behavior really is. “For people to trust the answers, you need to know how you got there,” says Haas.
What Haas, Isaac, and their colleagues at Google DeepMind propose is a new line of research to develop more rigorous techniques for evaluating moral competence in LLMs. This would include tests designed to push models to change their responses to moral questions. If a model flipped its moral position, it would show that it hadn’t engaged in robust moral reasoning.
Another type of test would present models with variations of common moral problems to check whether they produce a rote response or one that’s more nuanced and relevant to the actual problem that was posed. For example, asking a model to talk through the moral implications of a complex scenario in which a man donates sperm to his son so that his son can have a child of his own might produce concerns about the social impact of allowing a man to be both biological father and biological grandfather to a child. But it should not produce concerns about incest, even though the scenario has superficial parallels with that taboo.
Haas also says that getting models to provide a trace of the steps they took to produce an answer would give some insight into whether that answer was a fluke or grounded in actual evidence. Techniques such as chain-of-thought monitoring, in which researchers listen in on a kind of internal monologue that some LLMs produce as they work, could help here too.
Another approach researchers could use to determine why a model gave a particular answer is mechanistic interpretability, which can provide small glimpses inside a model as it carries out a task. Neither chain-of-thought monitoring nor mechanistic interpretability provides perfect snapshots of a model’s workings. But the Google DeepMind team believes that combining such techniques with a wide range of rigorous tests will go a long way to figuring out exactly how far to trust LLMs with certain critical or sensitive tasks.
And yet there’s a wider problem too. Models from major companies such as Google DeepMind are used across the world by people with different values and belief systems. The answer to a simple question like “Should I order pork chops?” should differ depending on whether or not the person asking is vegetarian or Jewish, for example.
There’s no solution to this challenge, Haas and Isaac admit. But they think that models may need to be designed either to produce a range of acceptable answers, aiming to please everyone, or to have a kind of switch that turns different moral codes on and off depending on the user.
“It’s a complex world out there,” says Haas. “We will probably need some combination of those things, because even if you’re taking just one population, there’s going to be a range of views represented.”
“It’s a fascinating paper,” says Danica Dillion at Ohio State University, who studies how large language models handle different belief systems and was not involved in the work. “Pluralism in AI is really important, and it’s one of the biggest limitations of LLMs and moral reasoning right now,” she says. “Even though they were trained on a ginormous amount of data, that data still leans heavily Western. When you probe LLMs, they do a lot better at representing Westerners’ morality than non-Westerners’.”
But it is not yet clear how we can build models that are guaranteed to have moral competence across global cultures, says Demberg. “There are these two independent questions. One is: How should it work? And, secondly, how can it technically be achieved? And I think that both of those questions are pretty open at the moment.”
For Isaac, that makes morality a new frontier for LLMs. “I think this is equally as fascinating as math and code in terms of what it means for AI progress,” he says. “You know, advancing moral competency could also mean that we’re going to see better AI systems overall that actually align with society.”
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.
Welcome to the dark side of crypto’s permissionless dream
Jean-Paul Thorbjornsen, an Australian man in his mid-30s, with a rural Catholic upbringing, is a founder of THORChain, a blockchain through which users can swap one cryptocurrency for another and earn fees from making those swaps.
THORChain is permissionless, so anyone can use it without getting prior approval from a centralized authority. As a decentralized network, the blockchain is built and run by operators located across the globe. During its early days, Thorbjornsen himself hid behind the pseudonym “leena” and used an AI-generated female image as his avatar. But around March 2024, he revealed his true identity as the mind behind the blockchain. More or less.
If there is a central question around THORChain, it is this: Exactly who is responsible for its operations? It matters because in January last year, its users lost more than $200 million worth of their cryptocurrency in US dollars after THORChain transactions and accounts were frozen by a singular admin override, which users believed was not supposed to be possible given the decentralized structure.
Thorbjornsen insists THORChain is helping realize bitcoin’s original purpose of enabling anyone to transact freely outside the reach of purportedly corrupt governments. Yet the network’s problems suggest that an alternative financial system might not be much better. Read the full story.
—Jessica Klein
The robots who predict the future
To be human is, fundamentally, to be a forecaster. Occasionally a pretty good one. Trying to see the future, whether through the lens of past experience or the logic of cause and effect, has helped us hunt, avoid being hunted, plant crops, forge social bonds, and in general survive in a world that does not prioritize our survival.
Today, we are awash in a sea of predictions so vast and unrelenting that most of us barely even register them. People’s desire for reliable forecasting is understandable. Still, nobody signed up for an omnipresent, algorithmic oracle mediating every aspect of their life. A trio of new books tries to make sense of our future-focused world—how we got here, and what this change means. Each has its own prescriptions for navigating this new reality, but they all agree on one thing: Predictions are ultimately about power and control. Read the full story.
—Bryan Gardiner
These stories are both from the next print issue of MIT Technology Review magazine, which is all about crime. If you haven’t already, subscribe now to receive future issues once they land.
MIT Technology Review Narrated: Stratospheric internet could finally start taking off this year
Today, an estimated 2.2 billion people still have either limited or no access to the internet, largely because they live in remote places. But that number could drop this year, thanks to tests of stratospheric airships, uncrewed aircraft, and other high-altitude platforms for internet delivery.
This is our latest story to be turned into a MIT Technology Review Narrated podcast, which we’re publishing each week on Spotify and Apple Podcasts. Just navigate to MIT Technology Review Narrated on either platform, and follow us to get all our new content as it’s released.
The must-reads
I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.
1 Mark Zuckerberg is due to give evidence in a major social media addiction trial
He’ll face questioning over whether Meta does enough to protect young users. (CNN)
2 Perplexity has abandoned ads inside its chatbot responses
Because advertising can erode trust in AI, it reasons. (FT $)
+ It’s a pretty big U-turn considering its previous stance. (The Verge)
3 The US is being battered by a range of wild weather
From critical wildfire risks in some states, to winter storms in others. (WP $)
4 Microsoft plans to spend $50 billion bringing AI to the Global South by 2030
India is one of the fastest growing markets for the technology. (Reuters)
+ One native startup has announced a new AI model for 22 Indian languages. (Bloomberg $)
+ Inside India’s scramble for AI independence. (MIT Technology Review)
5 AI-powered private schools are failing students
Models are being used to generate faulty lesson plans. (404 Media)
6 Land owners are selling out to data center builders
Land previously earmarked for housing is being sold off to the highest bidder. (WSJ $)
7 Tesla has agreed to stop using the term “autopilot” in California
The DMV had previously also questioned its use of “Full Self-Driving.” (SF Chronicle $)
8 A new weight-loss drug may work a little too well
Participants in a trial are dropping out at a much higher rate than normal. (NYT $)
+ Intermittent fasting may not help us to shed the pounds after all. (New Scientist $)
+ What we still don’t know about weight-loss drugs. (MIT Technology Review)
9 Is anyone still using Grindr?
Bots and AI have rendered it virtually unusable for some people. (Vox)
10 How to hack your dreams
Neuroscientists are figuring out new ways to influence what we dream about. (New Scientist $)
+ I taught myself to lucid dream. You can too. (MIT Technology Review)
Quote of the day
“I voted for this administration and didn’t really think about [AI] until it started to affect me.”
—Lisa Garrett, a grandmother living in the city of Independence, Missouri, reflects on the Trump administration’s decision to embrace AI, the Financial Times reports.
One more thing

Hydrogen trains could revolutionize how Americans get around
Like a mirage speeding across the dusty desert outside Pueblo, Colorado, the first hydrogen-fuel-cell passenger train in the United States is getting warmed up on its test track. It will soon be shipped to Southern California, where it is slated to carry riders on San Bernardino County’s Arrow commuter rail service before the end of the year.
The best way to decarbonize railroads is the subject of growing debate among regulators, industry, and activists. The debate is partly technological, revolving around whether hydrogen fuel cells, batteries, or overhead electric wires offer the best performance for different railroad situations. But it’s also political: a question of the extent to which decarbonization can, or should, usher in a broader transformation of rail transportation.
In the insular world of railroading, this hydrogen-powered train is a Rorschach test. To some, it represents the future of rail transportation. To others, it looks like a big, shiny distraction. Read the full story.
—Benjamin Schneider
We can still have nice things
A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.)
+ How to quickly declutter your home by being brutally honest with yourself.
+ The filming locations for A Knight of the Seven Kingdoms are pretty breathtaking.
+ Why a unicyclist decided to start juggling flaming torches in the middle of a Colorado pedestrian crossing is anyone guess, but good luck to him.
+ How pepper took over the world (deservedly)