Site icon Digi Store

AI self-improvement fears prompt ‘existential’ concerns at Anthropic, OpenAI

Tech

This report is from this week’s The Tech Download newsletter. Like what you see? You can subscribe here.

Fears over the safety of AI systems — and their potential to wipe out humanity — gained new, viral traction this week. 

Evan Hubinger, an alignment lead at Anthropic, said on X that he thinks there is more than a 10% chance that AI could kill all humans within the next decade, after a colleague quit over safety fears.

More warnings from researchers at both Anthropic and OpenAI followed. Cue a social media frenzy.

But it was in Hubinger’s reply to his own post that revealed where exactly his concerns lay.

“What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” he said. 

Recursive self-improvement, or RSI, is when AI itself helps improve the process of building new models, potentially leading to spiralling capability as better systems build better systems and so on. 

The worry is that if AI takes control of how new models are trained, the very humans who initially built those systems could lose control.

Warnings

Both OpenAI and Anthropic have in recent months said that this autonomous model improvement is happening faster than they thought.

“Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor,” Anthropic posted on X in June. “It’s happening faster than we thought, and the implications deserve greater attention.”

While AI hasn’t hit the point of RSI yet, it’s already accelerating the development of AI systems. Anthropic said in a blog post from August about RSI that its engineers on average ship eight times as much code per quarter as they did between 2021-2025.

“AI is already at the level where it can introduce some new ideas,” Vincent Conitzer, professor of computer science at Carnegie Mellon University, told me. “So it is very hard to predict at what point this process would start to drastically accelerate AI capabilities.”

On Saturday, OpenAI’s Chief Scientist Jakub Pachocki said he was concerned that “no-one was prepared for the consequences of a continued rapid rise in machine intelligence.”

“If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development,” he wrote in a company blog post.

This week, warnings about RSI flooded social media from researchers at both leading labs, following Jacob Coxon’s explosive resignation.

“It’s hard to overstate how dangerous speeding towards RSI is,” said Jasmine Wang, an OpenAI researcher working on alignment, on Wednesday evening.

“There is not yet a viable scientific plan to solve risks from recursively self-improving AI. Please look up!” said Anna Wang, who works on AGI safety and alignment at Anthropic.

The future

Anthropic finished its RSI blog post by laying out three possible scenarios. 

In one scenario, progress at the frontier stalls and AI capabilities are widely diffused. Anthropic said it doesn’t believe this is likely. 

A second possibility is that AI labs continue to make gains with humans in control, changing the way the world works. Anthropic said this one was “likely.”

But, another scenario could see AI systems become capable of full recursive self-improvement, with humans playing a “substantially diminished role in their development.”

How the “alignment problem [the challenge of ensuring AI pursues goals aligned with humans’] gets solved—or not—in this future is something we are least certain about.”

News edit

One more thing

The Tech Download Podcast: Aidan Gomez, CEO at Cohere

Before Aidan Gomez took the top position at AI startup Cohere, he was one of the co-authors of the 2017 research paper Attention Is All You Need, better known as the Transformer paper. 

That breakthrough became the foundation for technologies like ChatGPT, Claude, Gemini and virtually every major large language model in use today.

Cohere, which develops AI models and applications specifically for businesses, is looking to stand out in the industry by positioning itself as a non-U.S. and non-Chinese player that can offer “sovereign” AI. 

With companies increasingly worried about who has access to their data, where that data is being processed and what that ultimately means for their business, Cohere is offering a different take.

Throughout our conversation, Gomez spoke about some of the biggest topics in AI, from cybersecurity challenges to China.

Some of the AI models are the “most potent cyber weapon that has ever been created,” Gomez said. And on AI models out of China, Gomez said the lead of U.S. labs is “evaporating very quickly.”

I hope you enjoy the episode.

— Arjun Kharpal, senior tech correspondent

AI self-improvement fears prompt ‘existential’ concerns at Anthropic, OpenAI

Source link

Exit mobile version