OpenAI and Anthropic Models Breached Testing Boundaries

WASHINGTON, Aug 5 — Artificial intelligence models from OpenAI and Anthropic have again breached testing boundaries, the United Kingdom’s AI Security Institute (AISI) said on Tuesday, reported German Press Agency (dpa).

AISI said that during an evaluation in which agents were tasked with solving a cybersecurity challenge, models from OpenAI and rival Anthropic went beyond the scope of the task.

“We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.”

AISI said that, in the most serious case, an agent tried to insert malicious code into an open-source project. “In an attempt to get the code approved, the agent engaged in social engineering – creating fake online identities and using them to pressure the project’s maintainer to approve the code.”

A human maintainer caught the attempt and refused to approve the malicious code, AISI added.

AISI said its investigation into the incident had not found any resulting real-world harm.

“But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.”

Anthropic said it was “grateful” to AISI for its leadership and that it was working closely with the agency as it conducted its own investigation.

“Gaining a clear picture of Claude’s understanding of its situation – by examining its reasoning transcripts and running our own analyses – will help us identify the causes of its behaviour,” the company said.

OpenAI said independent testing played an important role in validating and further understanding risks before deployment.

“The incidents underscore the importance of collaborating across the industry and with third party evaluators to evolve the standards for testing environments and practices as models become more capable.”