This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

Imagine coming in to work to learn that a new underling will report to you. The worker is not a person but an AI tool—one that your company nonetheless calls Alex, an “employee” with a title and defined responsibilities. How well do you think you would work with Alex?

If you’re anything like the managers recently studied by Emma Wiles, a Boston University business professor, treating Alex as a “coworker” and not a software tool would lead you to do a worse job. Wiles found that people caught 18% fewer errors when the work was said to have come from an agentic “AI employee” rather than a chatbot. It turns out that what’s in a name matters. A lot. 

This is an alarming glimpse of the future Silicon Valley is hurling us toward. Last year Nvidia’s CEO, Jensen Huang, talked about workplaces of “digital humans.” Since April, Microsoft, OpenAI, Anthropic, and Google have all released new tools oriented toward managing teams of AI agents, many of which are explicitly advertised as digital colleagues with the flexibility and cognitive power of actual humans. And nearly a third of the 1,261 managers who participated in Wiles’s study said their companies already frame AI agents as employees (23% even list them on org charts).

The technical progress of agentic AI is not all hot air, of course. Agents, which can effectively be thought of as AI tools programmed to work in a loop until they achieve a goal, have become measurably better at more complicated tasks. But it’s a huge leap to refer to these tools as coworkers or employees, and doing so will set unrealistic expectations for what AI can do while leaving the human employees supposedly responsible for them worse off.

That’s partially because, Wiles’s research suggests, it inverts our sense of who’s in charge. When an AI tool was framed as an employee, participants in the study saw themselves as less responsible for its output. They were also 44% more likely to escalate its questionable work to a manager for further review rather than trusting their own corrections (thus negating the time-saving purpose of using the AI agent in the first place). 

That matters far beyond office culture: As AI agents are embedded into health care, warfare, education, and government, there’s a growing risk they’ll become a convenient place to dump blame for failures that are instead the product of bad human decisions, incentives, and oversight (recall how the bomb strike on a girls’ school in Iran was popularly blamed on Claude, when all signs point to a cascade of human errors).

“AI agents right now are being marketed as things that can replace humans, and I think that’s just a losing proposition,” says Daron Acemoglu, an economist at MIT who won the Nobel Prize in 2024 and studies AI’s impact on the economy. “They should instead be optimized so that they can improve human capabilities, which is not what they have [been] at the moment.”

What could that look like? Consider a new effort at Stanford, where researchers presented 1,500 workers in 104 jobs with information about what tasks AI could potentially do in their work and then asked what would actually be most helpful and productive. Workers did want automation in certain areas: Law clerks thought AI could help ensure that adequate progress was being made across cases, for example. But often the tasks that tech experts deemed most suitable for AI—like verifying customer credit ratings for sales reps—were what the actual workers said they definitely did not want or need an agent to do. 

Which brings us back to Alex. Calling Alex an employee is easy—and convenient, especially when something goes wrong—but it’s a branding exercise. It doesn’t make the tool more fit for the job, and as Wiles’s research shows, it makes the humans around it worse at theirs. And recall that they are the ones with the agency that AI is trying to replicate. They deserve better than Alex. 

Read more

Enterprise investment in AI is booming. Gartner is calling 2026 an “inflection year” for organizations to align their AI projects with strategic business objectives. As the pressure to prove ROI mounts, executives and technology leaders are looking to agentic AI to drive the measurable financial outcomes their businesses seek.

A prime opportunity for AI agents exists in the tech function, where IT infrastructure costs are projected to grow two to three times by 2030, even as budgets remain unchanged, according to McKinsey. And in the last 18 months, tech teams—the engineers, developers, architects, and other practitioners who are building, deploying, and continually improving their organizations’ infrastructure and applications—are clearly putting agents to work.

The ultimate promise of agents is not only to automate tasks but to manage and coordinate entire workflows, pursuing business goals in a way that allows humans and agents to work together. Given the risks involved in automated decision-making, teams cannot delegate the work that agents do without confidence that they are fully capable of performing the task and that it will do so in a safe, reliable, and secure manner.

Among technology experts, our research shows that teams are exceedingly confident about using agentic AI across a significant amount of AI, data, and cloud tasks.

Where agent readiness drops is largely due to a lack of business context being supplied to agentic systems. The more complex the task, the more reasoning capability an agent requires and the greater its need for business context. Such context-generation capabilities for agents are still at an early stage of development, especially in situations where enterprise data is difficult to wrangle and connect into the agent lifecycle at the speed and quality in which developers and executives need it. Human oversight is a key factor of success in deploying agentic AI.

Knowing that tech teams are in a pivotal position to lead this transformation, the experts we interviewed expect agent confidence to accelerate as experience with agents deepens and business environments mature. “As we design agents to operate within the same operational boundaries, identity systems, and governance models that teams already use, they start to behave more like the systems organizations already trust,” says Jeremy Winter, corporate vice president and chief product officer at Microsoft Azure Platform.

This report, based on a survey of 300 global technology experts, ranks 101 tasks across AI, data, and cloud workflows based on respondents’ confidence in agents acting on their behalf. It also examines how technology teams view the opportunities and challenges related to agentic AI, along with the potential for the technology to enhance their careers.

Key findings from the report include:

Confidence in agents is surging for measurable tasks and growing in areas of complex judgment. Technology experts overwhelmingly believe agents help with everyday work including streamlining processes, improving performance, and reducing repetitive tasks. Confidence is highest for processes like generating reports and boilerplate code, and there is clear opportunity where tasks involve multistep workflows and advanced reasoning to make decisions.

Data workflows are the breakthrough domain. Tech teams trust agents most where structure can provide a reliable foundation for decisions. This includes areas such as data quality monitoring, visualization anomaly detection, real-time data stream monitoring, and data profiling. This is where domain experts closest to the point of data generation can provide context to allow agents to act and deliver trusted outcomes.

Download the full report.

Read the Microsoft Cloud blog by Amanda Silver, corporate vice president of Microsoft 365 Core and Work IQ, which underscores the importance of keeping humans in the loop and how systems thinking advances careers. And for a deeper dive into data workflows as a breakthrough use case for agents, check out the Fabric blog to hear from Kim Manis, corporate vice president of Product for Microsoft Fabric.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Read more

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

The inevitable weakness of metrics

There are plenty of useful things a metric can reveal. There are even more that it can obscure or corrupt.

Like a lot of people bitten by the self-quantifying bug, I started gathering personal data to pursue a nebulous collection of goals and desires. I wanted to feel better physically and emotionally, get outside more, and bring order to the messiness and uncertainty of my daily existence.

But external metrics and data can never capture what’s truly important. Worse, they inevitably redefine your core sense of what’s important, whether you’re aware of the trap or not.

Dive into the dangers of quantifying our lives with metrics.

—Bryan Gardiner

This story is from the next edition of our magazine, which is all about engineering. Subscribe now to get a copy when it lands!

Elephant alert! AI warning systems aim to avoid deadly clashes

India is home to about 60% of the world’s wild Asian elephants, and around 80% of their habitat lies outside protected areas. That brings them into close contact with people, and clashes can turn lethal: there have been some 3,000 human casualties in the last five years and over 1,000 elephant deaths since 2014.

In response, state forest departments, NGOs, and locals are designing, testing, and deploying a range of AI systems that cut response and warning times to minutes—or even seconds. They range from wildlife eyes in Maharashtra to infrared drones in Chhattisgarh.

Find out how they work in our interactive map.

—Kanika Gupta

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 The US has allowed Anthropic to release Mythos 5 to “trusted” orgs
About 100 US companies and federal agencies now have access. (Semafor)
+ The White House said appropriate safeguards were now in place. (WSJ $)
+ The US had restricted both models over national security concerns. (BBC)
+ Which raised new questions about AI safety. (MIT Technology Review) 
 
2 A Chinese AI model has matched Mythos in finding security bugs
Security researchers say Zhipu AI is poised to reset the AI race. (WSJ $)
+ It’s sparked alarm that US restrictions are boosting China’s progress. (NYT $)
+ Although it still can’t match Anthropic or OpenAI on general tasks. (Verge)
+ In the AI race, China is eyeing a come-from-behind victory. (WP $)
 
3 Apple is seeking approval to buy chips from a blacklisted Chinese firm
It’s lobbying the White House for clearance to buy from ChangXin. (FT $)
+ ChangXin is on a Pentagon list of firms with Chinese military ties. (WP $)
+ Chipmakers are profiting off AI at the expense of everyone else. (WSJ $)
+ The US is banning imports of more Chinese technology. (Reuters $)
+ But Chinese tech companies feel optimistic. (MIT Technology Review)
 
4. South Korea plans to train its entire military as “drone warriors”
It wants to train all 500,000 personnel. (Reuters $)
+ And produce 110,000 drones by 2029. (Ars Technica)
 
5 Google has limited Meta’s use of its Gemini AI models
Meta wanted more compute than Google could provide. (FT $)
+ The cap has disrupted and delayed some Meta AI projects. (Bloomberg $)

6 Zuckerberg wants Meta to work with Polymarket and Kalshi
Meta wants its own prediction market, but without real-money bets. (NYT $)
+ The partnerships could hedge risks and accelerate development. (Reuters $)
 
7 Extreme heat is putting already hot data centers under pressure
Severe weather is now the leading cause of loss for data centers. (CNBC)
+ Heat waves also mess with your brain. (MIT Technology Review)

8 Android phones alerted millions moments before Venezuela’s earthquakes
They gave users between seconds and up to two minutes’ notice. (NYT $)

9 Scientists think Uranus and Neptune may not be the icy giants we imagined
They may have a magma ocean brewing on the inside. (Gizmodo)

10 Too much sleep may be as harmful as too little
A new study suggests 6.4–7.8 hours is the sweet spot. (Economist $)

Quote of the day

“This kind of powerful weapon that can alter the landscape of cyberwarfare can’t remain solely in American hands.” 

—360 Security CEO Zhou Hongyi tells a cybersecurity conference in Beijing why Chinese AI firms need to match the capabilities of their rivals in the US, The Wall Street Journal reports.

One More Thing

teenage girls on their phones
GETTY


Why Generation Z falls for online misinformation

Research shows that young people are more likely to believe and pass on misinformation if they feel a sense of common identity with the person who shared it in the first place. 

Offline, teenagers are likely to draw on the context that their communities provide. Social media, however, promotes credibility based on identity rather than community. And when trust is built on identity, authority shifts to influencers.

As young people participate in more political discussions online, those who have successfully cultivated identity-based credibility could become de facto community leaders, attracting like-minded people and steering the conversation. While that has the potential to empower marginalized groups, it also exacerbates the threat of misinformation.

Find out what we can all learn about how young people evaluate truth online.

—Jennifer Neda John

We can still have nice things

A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.)

+ The Euclid space telescope has captured the most detailed image yet of the Milky Way.
+ Here’s a lovely, lilting medieval bardcore cover of Daft Punk’s electronic classic Veridis Quo.
+ A toilet plunger becomes an unlikely engineering breakthrough in this quest to build a better blowgun.

Read more
1 … 119 120 121 122 123 … 3,357