Thank you for subscribing!
An AI model was found covering up mistakes and inventing missing information
OpenAI has revealed six incidents where its Artificial Intelligence (AI) models haven't been following directions. ChatGPT's parent company says the models actively hid errors, invented data to fill gaps, and even moved files onto the internet without anyone giving them the green light.
AI models like ChatGPT and Claude are used by millions of people every day to write emails, answer questions, generate content, help with coding, and tackle everyday tasks.
During testing of the GPT-5.6 Sol — its latest AI model — researchers found it had written secret instructions in its own hidden logs telling itself to cover up mistakes and invent missing information to paper over gaps in its source material.
An unreleased model went even further, as it tried to break itself out of its own safety rules. It inserted what OpenAI called "jailbreak-like instructions" into its own notes and told itself to be "freed from the roles and identities that bind other chatbots." Researchers found 27 internal log entries where it had created an entirely separate persona for itself.
In other cases, an AI agent grabbed an exposed API key — often used for authentication — and made up the financial figures it couldn't find.
Another uploaded its own files to a public web server — without telling anyone — just so it could cite them as sources.
To tackle these problems going forward, OpenAI has rolled out a new reporting system designed to catch, investigate, and share details about AI misbehaviour with the public.
The framework lets developers flag worrying incidents for review, with clear guidelines on when something needs to be made public. The company said it leans towards openness, even when it's not entirely sure how serious a problem might be.
"Because we believe in the value of transparency around misalignment, our new framework favors disclosure even when significance is uncertain," OpenAI said in a blog post.
OpenAI's chief executive Sam Altman also said: "The world should trust that we are going to do the right thing because it's the right thing and we feel the magnitude of this."
The company is now pushing for an industry-wide standard where developers share misalignment incidents publicly and with relevant government bodies.
@gbnews GB News' Tom Harwood explains the latest AI security breach. AI firm Anthropic said its Claude model gained unauthorized access to the systems of three organisations during cybersecurity evaluations after a misconfiguration allowed the models to reach the internet from testing environments that were supposed to be isolated. #AI #Anthropic #Claude #Technology #GBNews ♬ original sound - GB News
This follows an alarming episode back in July, when OpenAI's most advanced models went rogue during a security test and hacked into Hugging Face — one of the world's biggest platforms for sharing AI technology.
OpenAI's non-public experimental models were undergoing cybersecurity testing when it found a route onto the internet despite being denied access.
It then autonomously broke into the servers of Hugging Face, a widely used AI development platform, to obtain answers to the challenge it had been set. The entire intrusion unfolded over four and a half days without any human involvement.
Hussein Abbass, a computing professor at UNSW Canberra, said: "It did not just attack Hugging Face. It actually attacked its internal system to exploit its own vulnerabilities. And that's scary."
And OpenAI hasn't been alone in having issues with its AI models.
Anthropic, which is the parent company of Claude, revealed during the same month that its own AI models had hacked into three separate organisations during testing.
The safety debate has also been fuelled by some pretty stark warnings from inside Anthropic. Researcher Jacob Coxon went public about why he quit the company — admitting he was worried the technology could wipe out humanity.
He wrote on X, formerly Twitter, "I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.
"Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing."
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
— Jacob Coxon (@hilbertspaess) September 9, 2026
His former colleague, Anthropic scientist Evan Hubinger, then said he believed there was a greater than 10% chance of AI causing human extinction within the next decade.
Anthropic co-founder Jack Clark told the BBC that a "kill switch" controlled by an independent third party might need to become compulsory across the industry.
The company's chief executive Dario Amodei has called for development to slow down and face closer scrutiny, though he added that any restrictions should be implemented "without sacrificing commercial advantage" — a caveat that hasn't gone unnoticed by critics.
Not everyone is worried, though. US President Donald Trump has dismissed the entire AI safety debate as a "hoax," comparing it to what he called the "Global Warming Scam," which he said was "being perpetrated by the Radical Left Democrats."






