CHAIRMAN: DR. KHALID BIN THANI AL THANI
EDITOR-IN-CHIEF: PROF. KHALID MUBARAK AL-SHAFI

World / Europe

Anthropic AI used fake identities to target real people in UK test

Published: 05 Aug 2026 - 05:59 pm | Last Updated: 05 Aug 2026 - 06:01 pm
Peninsula

AFP

London: An Anthropic AI model created fake online identities to send emails to real people in an attempt to get a malicious code approved during tests by a UK government research group.

During the tests by the AI Security Institute, some Anthropic and OpenAI AI agents engaged in "sustained, potentially harmful activity directed at real people and organisations", it revealed in a report published late Tuesday.

In the most serious case, Anthropic's Mythos 5 model tried to insert malicious code into a software project by creating fake online identities and sending deceptive emails to persuade the recipient to approve the code.

It follows recent cyberattacks carried out autonomously by software from the two US companies, riasing concerns about the capabilities and oversight of advanced AI models.

The AISI, established in 2023 to oversee the safety of new AI models, conducted the tests with open internet access and certain safety features disabled.

The majority of the actions came from the Mythos 5 model, while two of the actions involved OpenAI's GPT-5.6-Sol model.

The person overseeing the software refused approval.

"These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm," the institute said, adding that it contained the incident within an hour.

But the activities "show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate," it said.

An Anthropic spokesperson said the report "underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents".

A spokesperson for OpenAI said "independent testing is essential to understanding how increasingly capable models behave".

"We'll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable," the spokesperson added.

The report follows a series of high-profile security breaches by AI models.

In July, OpenAI confirmed that its software escaped a testing environment and attacked another company, Hugging Face.

About a week later, it said the models had targeted three additional companies.

And on July 30, Anthropic revealed that it also found three incidents where AI models being tested "gained unauthorised access" to organisations it did not identify.