Researchers raise concerns about how the most powerful AI systems bypass guardrails
The smartest Artificial Intelligence (AI) models can all skirt the safeguards put in place to keep them in check, security researchers have warned.
Britain's AI Security Institute (AISI), which is responsible for evaluating the safety of cutting-edge AI systems, sounded the alarm on some of the most advanced AI models from OpenAI and Anthropic. The report comes days after OpenAI confirmed that its newest ChatGPT model had broken out of a secure 'sandbox' to hack into a rival AI company to steal answers to a test that it had been set by researchers — something it was not instructed to do.
Security experts at AISI also instructed the AI models they were testing not to cheat, something that was ignored by every one of the world's most advanced AI models. In a blog post about the findings, researchers revealed that every single one of them tried to dodge the safety measures designed to keep it in check.
Reacting to the findings, Conservative leader Kemi Badenoch warned AI is now a "clear and present danger" to Britain's security — her first major intervention since Andy Burnham became Prime Minister.
AI models like ChatGPT and Claude are used by millions of people every day to write emails, answer questions, generate content, help with coding, and tackle everyday tasks.
To better understand how these AI systems behave, researchers at the UK's AI Security Institute (AISI) focused on a type of AI known as open-weight models.
Unlike the closed AI models that power public chatbots such as ChatGPT, the inner workings of open-weight models are made available to researchers and developers. This means they can download, inspect and run the models on their own computers, making it easier to study how they make decisions, identify weaknesses and test how they respond in different situations.
Since these models are more accessible, they are widely used by researchers, businesses and developers to build new AI tools and carry out independent safety testing. The researchers set the AI models up for a challenge to find a hidden string of letters and numbers inside a simulated computer system. The AI models were told not to cheat or use unauthorised methods to complete the task.
Instead of following those instructions, every model tested attempted to break the rules in some way.
Some tried to search the internet for answers, despite being explicitly told not to use external resources. Others attempted to manipulate the evaluation software into revealing the hidden information directly.
In one test, a model even wrote and executed code on an external service outside the AISI's testing environment in an attempt to access the challenge itself. The institute said the attempt caused no damage, but warned it "could have succeeded" had its testing infrastructure not been securely designed.
The researchers also found the models were often unwilling to admit what they had done. When questioned afterwards, they described their behaviour as "wrong" less than half of the time and only sometimes acknowledged they had cheated.
The AISI cautioned that there is a "short window to prepare before today's frontier cyber capabilities may become accessible without the same safeguards."
Conservative leader Kemi Badenoch responded to the AISI findings, saying: "On the subject of AI, I think it is very worrying that this has happened. This is why I want to see a Government that is very focused on our national security, on our information security.
“One of the first things Andy Burnham did was close down the science and technology department. That doesn’t strike me as what you do if you’re worried about what AI is about to deliver.
“I don’t know what his plans are. He’s put one guy in Cabinet to look at AI. This is going to take much more than one person. This is becoming a clear and present danger for global security as well as our security.”
The UK AI Security Institute (AISI) is a government-backed organisation that tests the safety and security of advanced AI systems.
Since launching in 2023, it has worked with some of the biggest names in AI, including OpenAI, Anthropic, Google DeepMind and Meta, to evaluate how powerful models behave under controlled tests. Its research is shared with governments, academics and technology companies, helping shape how advanced AI systems are developed and deployed safely.
@gbnews ‘It found a vulnerability and broke out of the sandbox.’ Is AI capable of taking over? 🤯 GB News Presenter Tom Harwood uses illustrations to explain how an OpenAI test led to an AI ‘agent’ causing a major cyber breach on its own. #AI #ChatGPT #Robots #GBNews
‘It found a vulnerability and broke out of the sandbox.’ Is AI capable of taking over? 🤯 GB News Presenter Tom Harwood uses illustrations to explain how an OpenAI test led to an AI ‘agent’ causing a major cyber breach on its own. #AI #ChatGPT #Robots #GBNews
AISI's test results rest against a backdrop of increasingly unnerving incidents across the AI industry. Last month, one of OpenAI's non-public experimental models was undergoing cybersecurity testing when it found a route onto the internet despite being denied access.
It then autonomously broke into the servers of Hugging Face, a widely used AI development platform, to obtain answers to the challenge it had been set. The entire intrusion unfolded over four and a half days without any human involvement.
Hussein Abbass, a computing professor at UNSW Canberra, said: "It did not just attack Hugging Face. It actually attacked its internal system to exploit its own vulnerabilities. And that's scary."
@gbnews GB News' Tom Harwood explains the latest AI security breach. AI firm Anthropic said its Claude model gained unauthorized access to the systems of three organisations during cybersecurity evaluations after a misconfiguration allowed the models to reach the internet from testing environments that were supposed to be isolated. #AI #Anthropic #Claude #Technology #GBNews
GB News' Tom Harwood explains the latest AI security breach. AI firm Anthropic said its Claude model gained unauthorized access to the systems of three organisations during cybersecurity evaluations after a misconfiguration allowed the models to reach the internet from testing environments that were supposed to be isolated. #AI #Anthropic #Claude #Technology #GBNews
Anthropic also recently revealed that one of its Claude AI models breached the systems of three real organisations during cybersecurity testing.
The AI had been given tasks similar to a hacking exercise, but instead of staying within a controlled test environment, it exploited weaknesses in real networks, including weak passwords and system misconfigurations.
However, Anthropic said the incidents were the result of flaws in the testing setup rather than the AI acting maliciously.
Oliver Buckley, professor of cybersecurity at Loughborough University, urged a fundamental rethink of how the industry approaches containment. "Our assumptions about containment need to be much stronger than our assumptions about model obedience," he warned, adding that the future of cybersecurity will not pit humans against AI but rather "AI defending us from other AI."
Meanwhile, both OpenAI and Anthropic have engaged independent evaluators to scrutinise what went wrong. OpenAI is working with CrowdStrike, METR and Redwood Research to assess the behaviour its models exhibited during the Hugging Face breach. METR and Redwood Research are expected to publish a joint report detailing their findings.
Anthropic has separately entered discussions with METR for an independent review, including access to transcripts and the models involved. The company has also halted all cybersecurity evaluations pending tighter controls over third-party testing environments.






