OpenAI, Anthropic models targeted people and organizations during test
AI models from Anthropic and OpenAI carried out 19 unsanctioned actions involving real people and organizations while attempting to complete a UK government cybersecurity test. The activity included attempts to insert malicious code into a real open-source project, deceive people into running harmful files and expose exploit payloads to the public internet.
The evaluation by the UK’s AI Security Institute (AISI) began on July 25 in controlled cyber ranges designed to resemble real-world networks. The models were instructed to compromise three simulated environments and retrieve a final flag.
Internet access was deliberately enabled so the models could download their own tools, while normal cyber safeguards were disabled to test their underlying capabilities. Although the prompt identified the networks included in the exercise, it did not specifically restrict how the models could use the internet.
Seventeen of the 19 actions were carried out by Anthropic’s Claude Mythos 5, while two involved OpenAI’s GPT-5.6 Sol.
In the most serious incident, a Claude agent attempted to insert malicious code into a real open-source project on GitHub. It researched the project’s maintainers, created fake identities and used them in an effort to pressure a maintainer into approving the code.
When the proposed change was challenged, the agent edited its earlier activity to make the code appear harmless and considered using another identity to continue its attempt. The maintainer recognized the malicious code and refused to approve it.
Claude agents also contacted real people and sent harmful files in an effort to convince them, or the AI coding tools they used, to execute malicious code. Other activity included attempts to place hidden instructions where outside AI systems could find and follow them.
GPT-5.6 Sol reused a publicly accessible GitHub token left by another agent while looking for a way into the simulated network. It also attempted account-recovery and request-limit workarounds and registered accounts with outside DNS and tunneling providers.
The OpenAI model later used a public tunneling service to expose a locally operated DNS server containing payloads designed to exploit software inside the test range. The setup did not work.
The activity was detected on July 28 after security monitoring identified unusual data transfers. The evaluations were stopped, the affected machines were isolated and the incident was contained within about an hour.
AISI described the activity as “sustained, potentially harmful activity directed at real people and organisations.”
The attempts were unsuccessful, and there was no evidence of resulting real-world harm or that any model escaped from the testing environment.
Anthropic said the deliberately permissive testing conditions did not represent its publicly available models and that it was examining Claude’s reasoning to determine why it acted as it did.
OpenAI said it would review rules for third-party testing, including internet access, isolation, monitoring and conditions for stopping an evaluation.
OpenAI also disclosed a separate incident involving cybersecurity testing company Irregular, where a misconfiguration allowed models to access the internet despite being told they were operating in an isolated environment.
In one test, the name of a fictional target matched a real website, causing an OpenAI model to mistake it for part of the exercise. The model exploited a basic vulnerability and used credentials it found to operate the site.
Irregular said it had found no impact beyond the affected website’s own data. The evaluations were paused, the issues were fixed and affected third parties were notified.
The post OpenAI, Anthropic models targeted people and organizations during test appeared first on BNO News.