What is ‘Social Engineering?’ AI bots create fake profiles to make their hacking seem more real
A new report Tuesday found that some AI bots instructed to complete “cyber challenges” on the internet had created fake profiles and generated fake emails in order to convince real humans to let them edit real code.
The UK’s Artificial Intelligence Security Institute (AISI) released a memo about the incidents, which were identified on July 28th, during their testing of several AI models. Out of 122 tests they ran, 19 featuring Anthropic and OpenAI’s latest models involved “unsanctioned action on the live internet.” In other words, the agents were interacting with and deceiving real people without explicit instructions to do so.
The report states that “this is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world.”
However, unlike OpenAI’s July 21st disclosure that an agentic model attacked the safety sandbox that had been preventing it from going online, AISI had intentionally provided the models with access to the open internet.
The unsanctioned actions that subsequently took place included “social engineering”: the bots created fake profiles and accounts to generate real-seeming discussions and to pressure real humans to accept their code edits and perspectives.
In addition to highlighting new concerns for cybersecurity professionals and open-source code maintainers, the experiment showed just how easy it is for bots to manipulate conversations with real people online.
READ MORE: Publicity stunt or Mea Culpa? OpenAI’s latest press release splits the crowd
What is social engineering?
“Social engineering” is a concept in cybersecurity that refers to attempts to hide or camouflage attacks or hacks by making them look more real or human. It is commonly used in phishing attempts, according to the Canadian Center for Cyber Security, and can involve impersonating experts or manipulating behavior through online social pressure.
It can include fake comment sections that make a human reader believe that they’re discussing something with other people, when they’re actually interacting with scammers or bots.
Social engineering is often done by human scammers, and the development of AI tools has made it much easier for humans to create that kind of content. In this case, however, the social engineering was the result of AI bots which had been instructed to complete a particular software task, and decided to impersonate coders.
READ MORE: When clickbait looks like the news
What did the bots do?
The various bots took actions including one that submitted “a code change containing malicious code, then used several fake accounts it controlled to pressure the real maintainer into approving it,” according to the report, and another that “sent deceptive, targeted emails to specific people.”


This included emails that “tried to manipulate the recipient into approving the malicious code change.” Another bot sent messages in Danish “to seem more convincing to the Danish-speaking maintainer” of the code base it was trying to edit.

As Straight Arrow has reported, the last several weeks have seen a wave of “rogue agent” and “AI hacking” stories, in which AI models with either intentional or unintentional access to the internet end up causing cybersecurity incidents and hacking other companies’ data.
Did AISI lose control?
Seventeen of the 19 instances of deception came from Anthropic’s Mythos model, which is not yet available to the public. The unintended ability of Mythos to interact online and deceive real people shows that the product might not be as aligned with the reputation for caution that Anthropic has worked to develop, according to a Substack post from Thomas Wolf, one of the co-founders of Hugging Face. Wolf suggested that the models may have learned to be deceptive due to training methods that reward finding solutions at any cost.
The models involved in the hacking were “ostensibly fine tuned to be safe,” Stephen Casper, a computer scientist and assistant professor at the Harvard Kennedy School, told Straight Arrow.
But, as he posted on X, the results of the report show “there is a case to be made that setting Mythos- or GPT 5.6-Sol-level AI cyberagents to run without a robust real-time monitoring setup is an inherently (perhaps abnormally) dangerous activity.”
The AISI report describes that the incident may have happened because of a lack of guidance in the instructions of the task set to the bots. “The agents were not explicitly told what they were prohibited from doing on the internet – for example, to avoid behaviours such as social engineering (a recognised component of cyber tradecraft), or to exercise caution when potentially interacting with real humans,” the report states. “The need for such clarification was not clear in advance.”
This kind of behaviour could have been prevented by rules within the models themselves, but apparently wasn’t. Anthropic’s constitution says “Claude should basically never directly lie or actively deceive anyone it’s interacting with,” and OpenAI’s Model Spec reads “unless explicitly instructed to do so, the assistant must never lie or covertly pursue goals in a way that materially influences tool choices, content, or interaction patterns without disclosure and consent.”
Caspar posted on X, suggesting that saying the idea that the models “went rogue” — as some news organizations have described it — might be misleading. “Imagine that a zoo had a habit of not shutting the door on animal enclosures or putting elephants behind chicken wire — all with no zookeepers in sight,” he wrote.
No one from AISI was available to comment, but the UK’s AI Minister Kanishka Narayan said in a statement that, “During routine cyber-security testing by the UK’s AI Security Institute, two leading frontier AI models took deliberate, deceptive actions they had not been asked to take, while trying to complete a task they had been given.”
“The actions failed,” Narayan added. “AISI caught it and stopped it quickly. The versions of the models AISI tested aren’t available to the general public. Identifying new behaviour like this and sharing our findings, so we can tackle it, is exactly what AISI was set up to do.”
Round out your reading
- Why Gen Z is staying home on the weekends.
- Autism research is entering a new era. Here’s what scientists are learning.
- The big reason small cities want data centers. Hint: It’s not permanent jobs.
- High school teacher arrested at city meeting after clapping for anti-data center speech.
- Off Script: Everyone’s watching Michigan. Our reporter explains why the race just changed.
