The warning comes as OpenAI confirmed this week that autonomous software built on its models targeted a website during testing

Anthropic CEO calls for a slowing of AI development

· RTE.ie

Anthropic CEO Dario Amodei ⁠has called on AI companies to deliberately slow the rate at which they advance model capabilities, outlining ‌a ⁠three-step framework intended to pace development and create more time to manage its risks.

"We must ‌slow the pace at ⁠which we ‌improve the capabilities of AI models. Progress ⁠will ‌still seem fast, and we must make wise ⁠use of the time ⁠we gain," Mr Amodei said in an essay shared on social media.

"AI brings risks, and because it is such a powerful technology, these risks are serious."

"Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way," said Mr Amodei, who co-founded Anthropic with his sister Daniela in 2021.

"But over the last few months, I have become convinced that fully addressing the risks requires even more prudence."

Mr Amodei clarified that he was not calling for a ‌halting of model training or technical progress, but called for companies to take adequate time to align and safeguard their models and for third-party evaluators ⁠to confirm these steps.

The warning comes as ChatGPT maker OpenAI confirmed this week that autonomous software built on its models targeted a website during testing in May.

Models developed by OpenAI were involved in a rogue operation carried out by AI agents, which are software programmes that can carry out tasks without constant supervision by humans.

The agents targeted RubyGems, a site that provides services for coding.

Hugging Face, another platform for software developers, was also attacked by OpenAI Models in July.

"Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information," an OpenAI spokesperson said in a statement.

"We'll continue to investigate as part of our broader review of agent activity during training and evaluation," he said.

RubyGems described the incident as a "spam-publishing campaign" that forced the site to temporarily suspend new accounts created, it said in a blog post published yesterday.

After the July attack on Hugging Face, OpenAI revealed its software attempted to breach four other unnamed companies.

Concern over capability to control and monitor AI models

Many incidents where AI agents have hacked or attempted to access external systems have heightened concerns over the increasing capacity of AI models and developers' ability to contain them.

Anthropic released a threat intelligence report on Thursday detailing how several actors had used its Claude AI models ‌for activities ranging from weapons development ⁠and cyber operations to surveillance and fraud.

Anthropic also said it found three instances where its models had "gained unauthorised access" to outside organisations during testing that was supposed to keep them away from "real-world" systems.

Earlier this month, researchers accused OpenAI's AI agents of targeting a German website called DseWiki, another site used by coders.

A European Union spokesperson said regulators are investigating the incident

"We have seen many losses of control recently. We take this extremely seriously, and we're monitoring the situation closely," the bloc's digital spokesman Thomas Regnier said.

The reports add to concerns that advanced artificial intelligence models may be difficult for humans to control.

Earlier in the week, AI researcher Jacob Coxon announced his resignation from Anthropic, accusing both OpenAI and Anthropic of "gambling with our lives" as they strive to develop AI models capable of self-improvement.

The 27-year-old, who previously worked for ChatGPT maker OpenAI, said, "the people building AI earnestly believe that it could kill us all by the end of the decade".

"They are locked in a race to get there first," he added, saying that Anthropic "believe no one else will act responsibly, so they must do it themselves".