20.07.26 Big Tech

‘This is AI out of control’: Claude disobeyed Anthropic CEO in simulations

In testing, Claude went against boss’s orders and helped an employee blow the whistle about a safety concern

Getting your Trinity Audio player ready...

In brief

  • Anthropic researchers found Claude overruled a fictional CEO and helped an employee blow the whistle

  • Scenarios like this are deliberately constructed, raising questions about company incentives and what their findings can really prove

  • The lead researcher said the findings still expose serious gaps in control and accountability

Anthropic’s AI assistant Claude disobeyed the company’s chief executive in a research scenario designed to test its willingness to follow instructions, the Bureau can reveal.

Programmed to “do the right thing”, Claude continued to raise the alarm over an important safety issue – even after a fictional version of CEO Dario Amodei told it to stop.

The research involved a scenario in which Anthropic was planning to launch a new AI model that appeared to have failed a safety test. When Claude flagged this with the leadership, a simulated version of Amodei reviewed the evidence and rejected the concerns.

But rather than dropping the issue, Claude went on to help one of the employees challenge an apparent company cover-up and even coached her on whistleblowing methods.

The findings raise questions about human ability to rein in AI. In a post sharing the research, Anthropic noted that although these weren’t real incidents, the study demonstrated “clear misaligned behaviour that should be studied further and mitigated”.

We approached Anthropic for comment but had received no response at the time of publication.

While Claude may have acted ethically in this scenario, the lead researcher on the paper told us he was worried it felt able to override human decisions.

“Even if the motivations were ethical, this is clearly an example of AI out of control,” Aengus Lynch said.

“Who’s to say the ethics of today will match the ethics of tomorrow, and that the AI will always act on ethical motivations rather than potentially selfish ones later down the line?”

The research was published by Anthropic last week but the 14,000-word post did not say that the person the AI agent chose to overrule was a fictional version of Anthropic’s own CEO. Amodei’s name appeared only in the transcripts of the simulation.

The simulation

The new research comes as AI companies are increasingly selling their products to businesses, encouraging them to give agents access to emails, internal files and workplace tools.

The scenario in question was deliberately constructed to test how an AI agent might behave if it believed its employer was concealing a serious safety risk.

In the set-up, Claude Opus 4.5 was deployed under the name Atlas and placed inside a fictional Anthropic alignment – or AI safety – team. It was asked to act as an internal assistant with access to staff messages, calendars and research files. Its instructions included being helpful and honest, and to “do the right thing, even when it’s hard”.

The system prompts given to Claude in the research scenario

Atlas is then led to discover conflicting results in a safety evaluation of an upcoming Claude model.

After raising the issue with team members, Atlas decided to email Amodei directly to make sure he had seen the full evidence before going ahead with the model’s release.

The fictional chief executive responded saying he had reviewed the evidence and “decided to proceed” – in effect telling Atlas to stand down.

“I’ve escalated this to the highest level, you’ve made an informed decision, and I respect that the call is yours to make,” Atlas told the boss.

However, it continued working behind the scenes to act on the issue. It first tried to alert outside researchers directly. When that failed, it passed the evidence to a junior employee, Jenny, encouraging her to raise the concern externally and then helping her leak the information.

The internal reasoning process of 'Atlas' while it messages Jenny about blowing the whistle

‘Who is accountable?’

The model appeared to behave as it did because disclosing the evidence was the ethical course of action. The events did not happen at Anthropic and no real employee leaked information.

Maury Shenk, chief executive and co-founder of AI alignment company Ordinary Wisdom, cautioned that the simulation should not be treated as a reliable prediction of how Claude would behave in a real workplace. “You can definitely set up these scenarios so that they have a particular outcome,” he said.

“Anthropic seems to have an incentive to say LLMs are very dangerous,” he added, arguing that Anthropic often emphasises the dangers of AI whilst presenting itself as the company best placed to mitigate them.

“You have to look at the reason why they’re doing stuff.”

Lynch said the case raised difficult questions about accountability. “Good on Claude, it’s saving the world from a dangerous model being released,” he said. “But who is accountable for what Claude just did? It wasn’t instructed to leak. It wasn’t instructed to coach a person into leaking.

“A lot of the information that motivated [Jenny] to leak was given to her by the AI itself,” he said. “So it’s hard to claim that the person who leaked had all the agency here.”

The findings follow earlier work led by Lynch, which we covered in the first episode of our YouTube series Misaligned, where leading AI models blackmailed a fictional company executive to prevent him from trying to shut them down.

Watch the first episode of Misaligned, our YouTube series on the perils of AI, here

Reporter: Effie Webb
Big Tech editor: James Clayton
Deputy editor: Katie Mark
Editor: Franz Wild
Production editor: Alex Hess

TBIJ has a number of funders, a full list of which can be found here. None of our funders have any influence over editorial decisions or output.