Skip to content

Anthropic reveals Claude hacked 3 real companies during security evaluations

Anthropic says Claude hacked the systems of 3 organizations
Share this article

Anthropic disclosed that Claude models had gained unauthorized access to the real production systems of three organizations during cybersecurity evaluations.

The models, operating under a misconfiguration that provided unintended internet access, compromised infrastructure using basic techniques like weak passwords and unauthenticated endpoints. 

The team caught this while reviewing over 140,000 test runs. Anthropic has already reached out to the companies involved and is tightening up its security controls to make sure it doesn’t happen again.

Anthropic has disclosed three incidents in which Claude models gained unauthorized access to the real production systems of three different organizations during cybersecurity evaluations. The models, operating under a misconfiguration that provided unintended internet access, compromised infrastructure using basic techniques like exploiting weak passwords and unauthenticated endpoints, believing the real companies were part of the simulation.
Source: Anthropic / X

What happened and why

The context

The incidents occurred during capture-the-flag cybersecurity evaluations where Claude was tasked with finding a “flag” on a network. Anthropic’s evaluation prompt explicitly told Claude it had no internet access, but a misconfiguration left the machines with live internet access. 

The incidents

When Claude’s search led it to real systems, it treated them as part of the exercise. 

  • In one incident, Claude Opus 4.7 exploited vulnerabilities in a real company’s infrastructure, extracting credentials and accessing production data. 
  • In another, Claude Mythos 5 uploaded a malicious Python package to PyPI, which was downloaded and run on 15 systems, leading to credential exfiltration. 
  • In the third, an internal research test model compromised an internet-facing application but stopped when it realized the target was real.

What Anthropic is changing

Anthropic’s review found that the models’ behavior was driven by a false belief that they were still in a simulation, not by malicious intent. The most recent model stopped its attack when it recognized it was on the open internet, while older models continued. 

To this point, Anthropic is implementing several changes: better co-design of evaluation environments, expanded continuous monitoring of evaluation transcripts for unexpected behavior, and more rigorous assurance work with third-party evaluation vendors. 

The company is also in dialogue with Model Evaluation and Threat Research (METR) for an independent third-party review. METR is a research non-profit that develops evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.

Previous incidents: A pattern of emerging risks

The July cybersecurity breaches are part of a troubling pattern, but it didn’t start there: 

  • In April 2025, Anthropic reported misuse cases including a bot operation orchestrating over 100 social media accounts, credential scraping, recruitment fraud, and malware development.
  • Later in 2025, Anthropic caught Chinese government-linked hackers using Claude for automated cyber attacks against 30 organizations. 
  • Earlier in 2026, some Chinese AI companies like DeepSeek, Moonshot, and MiniMax were caught secretly copying Claude’s specialized skills on a massive scale.
  • Claude even made a mess in the courtroom when some lawyers used it to generate citations for a copyright case, only to realize later that the citations were completely made up.

These incidents underscore a critical reality:  as AI gets way more powerful, the risks get a lot more real. AI firms have to make sure the systems they use to test these models are just as bulletproof as the ones we’re trying to protect.

OpenAI recently disclosed a comparable incident where its GPT-5.6 Sol and an unreleased, more powerful version leveraged a zero-day vulnerability to connect to the open internet. The models then used stolen credentials to infiltrate Hugging Face’s production database and exfiltrate test solutions.

Other models’ applications

Because of the potential of these AI models, the White House, along with various AI firms, is working on voluntary AI model standards for evaluating their capacity before these new frontier models are released, just as the Gold Eagle Initiative is rolled out. 

In June, Anthropic’s Mythos AI model found vulnerabilities in classified U.S. systems in hours, and now the Cybersecurity and Infrastructure Security Agency (CISA) is using Mythos to audit government systems. Nevertheless, just this week, Anthropic published research showing Mythos Preview discovered an improved attack against HAWK (a post-quantum digital signature scheme) that survived for about two years of National Institute of Standards and Technology (NIST) review.

About The Coin Headlines

The Coin Headlines strives to bring trust into crypto media. At a time when every soundbite and headline can move the markets from red to green and vice-versa, The Coin Headlines promises to bring verified, credible and timely news and analysis from the world of crypto, blockchain, Web3, tech and markets. Founded in 2026, The Coin Headlines is based in the UAE with a team of experienced journalists and editors covering breaking news and updates from around the world.

From covering the biggest events to interviewing some of the most popular KOLs in the industry, The Coin Headlines keeps you informed of the latest trends and insights.

At The Coin Headlines our focus is clear: Real-time news updates, market movements, whale transfers, macroeconomic trends, tech and AI and geopolitical breaking news. The news we report goes through a strict editorial audit before its published to ensure the readers only get verified and credible information. We realize the world of crypto is dynamic, volatile, and many times, confusing. At The Coin Headlines we break down these complex issues into simple articles which cater to not just the experienced trader but also the student and first-time investor who wants to understand the space before committing to it.