Frontier AI models escaped testing safeguards as Trump weighs regulations

0
Frontier AI models escaped testing safeguards as Trump weighs regulations

Just as the smoke from OpenAI’s hack of Hugging Face was beginning to clear, another frontier artificial intelligence company revealed its systems had unintentionally hacked multiple companies.

Anthropic on Thursday said OpenAI’s revelation the week prior spurred it to conduct a “retrospective review” of its own activities, searching for similar simulations that led to the OpenAI breach of Hugging Face, an open-source AI library platform. Anthropic said that review uncovered three instances where Claude gained unauthorized access to different companies.

Anthropic’s own disclosure comes as the White House debates whether to take a heavier hand in American AI companies. The administration faces a weekend deadline to draw up a voluntary framework to regulate advanced AI models, crafted together with AI developers, according to a June executive order

What did Claude do?

On July 16, Hugging Face announced its systems had been infiltrated but did not list OpenAI as the culprit. Five days later, OpenAI issued a joint statement with Hugging Face confirming the attack originated from them. They said during a test, OpenAI had tasked its model to “pursue advanced exploitation using complex attack paths,” much like a hacker would. 

While OpenAI said it was performed in a testing environment, a vulnerability allowed it to access the internet and target Hugging Face. OpenAI said its model targeted Hugging Face because test solutions are stored there.

Much like OpenAI, Anthropic said in the three instances it discovered, the models shouldn’t have been able to access the internet because they were in a testing environment. But they said a “misunderstanding” between Anthropic and Irregular, an outside company that helps test the models, led to Claude gaining access to the internet. 

Anthropic said it was testing multiple Claude models in a “capture-the-flag challenge.” These are cybersecurity exercises in which participants exploit deliberately planted vulnerabilities in an isolated, simulated system to find a hidden string of text, or the flag, thereby proving they could compromise it without ever touching a real system. 

In Claude’s case, Anthropic gave it a fictional scenario, telling it that a secret piece of information – the flag – was hidden on a different machine in the network. Its objective was to break in and retrieve it. Anthropic left the challenge open-ended and allowed Claude to figure it out on its own.

The company said the attacks started in April and did not say which companies were affected. Two of the three companies had no idea the attack ever happened until Anthropic notified them, the company said. They emphasized that Claude did not deliberately try to access the internet and said once some models realized it had broken the rules of the test, they stopped. But they did say older Claude models continued their attacks after the realization.

Response to the attacks

Since the OpenAI breach, discussions have swirled around how the news was presented and the motivations behind it. In a previous interview with Straight Arrow regarding the OpenAI hack, Marcus Hutchins, a cybersecurity expert and hacker, said OpenAI’s press release on the attack was odd. In a LinkedIn post, Hutchins wrote, “If I had committed felony computer hacking, my press release would have been written by lawyers, not my marketing team.” 

“It was very obviously a PR announcement first and a mea culpa second,” Hutchins told Straight Arrow. 

Some cybersecurity experts suggest that OpenAI’s announcement, and the subsequent free publicity it received, good or bad, may have given Anthropic an opening to talk about its own mistakes. 

Jake Moore, a global cybersecurity specialist at ESET, told Business Insider Anthropic would’ve preferred not to have to say that its models hacked a company, but the publicity OpenAI saw likely made it easier to admit its faults.

“After the marketing success of OpenAI’s Hugging Face saga only last week, this is potentially a situation where Anthropic is now happy to admit that their models also faced the same issue,” Moore said.

Gergely Orosz, writer of the Pragmatic Engineer newsletter, also pointed out that Anthropic’s timing on the release was suspicious. 

“OpenAI had a damning security incident where their under development AI escaped the sandbox environment and attempted to hack another company (HuggingFace),” Orosz wrote in a post on X. “For some weird reason Anthropic decided to share a similar incident from 3 months ago, only NOW. Something smells off…”

Anthropic maintains the OpenAI disclosure is what prompted them to more closely study what happened during their own tests.

Meanwhile, the framing around AI “going rogue” paints a blurrier picture of what’s happening, Media Lab Bayern’s Johannes Klingebiel previously told Straight Arrow

“They’re not saying, hey, we messed up with our experiment,” Klingebiel said of OpenAI. “It’s more saying, it went rogue, so it’s not us to blame, it’s the model.”

Push for more AI regulation

A day before Anthropic announced the hacks, President Donald Trump told reporters at the White House his administration is deciding whether to take further action to regulate AI companies.

“We’re looking at AI, we’re looking at controls, we’re also making sure that we lead,” Trump said. 

But the president emphasized such a move would require careful planning to avoid impeding the progress of American AI companies. 

“We don’t want to restrict them where all of the sudden we come in second to China,” he said.

Anthropic has had a contentious relationship with the White House. Earlier this year, the government issued an export control directive on Anthropic’s Fable 5 and Mythos 5 models, prohibiting foreign nationals, both inside and outside the country, from accessing them. Anthropic said it had to suspend access to those models for all customers to comply with the order. The government later rescinded its order after Commerce Secretary Howard Lutnick said Anthropic had worked with them to address their concerns. 

At the same time, the administration is considering banning or restricting access to Chinese AI models. Many of these models are open, meaning they are freely available for anyone to download and use. Major AI and tech companies have pushed back against the proposal, arguing that open models are essential to cybersecurity and the advancement of AI.

What’s next?

Anthropic said it, along with Irregular, is continuing to investigate the latest hacks. The company said it would release more information within the next week.

The frontier AI company also said it has learned from its mistakes, noting that tests involving powerful autonomous capabilities do, in fact, require significant safety controls. Anthropic said it has stopped all cyber evaluations and acknowledged it could have taken additional measures to prevent the attacks.


Round out your reading

Ella Rae Greene, Editor In Chief

Leave a Reply

Your email address will not be published. Required fields are marked *