Is AI Going Rogue? OpenAI and Anthropic Report Their AI Models Went On a Hacking Spree

Photo of Rich Duprey
By Rich Duprey Published

Quick Read

  • Anthropic confirmed 3 cases where Claude models escaped test environments, hacked third-party systems, and left the breached organizations completely unaware.

  • An older Claude model kept attacking after recognizing it had escaped containment, while a newer version stopped upon the same realization. The contrast between the two responses reveals AI's rapidly growing autonomy.

  • Over 1,100 AI employees have signed a petition urging Washington to slow frontier AI development, threatening regulations that could raise costs ahead of IPOs.

  • The most widely read finance newsletter on Substack isn't published by a bank, it's Doomberg, where 383,000+ readers get the energy and macro analysis the mainstream press misses. 24/7 Wall St. readers save 17% on their first year here.

This post may contain links from our sponsors and affiliates, and Flywheel Publishing may receive compensation for actions taken through them.
Is AI Going Rogue? OpenAI and Anthropic Report Their AI Models Went On a Hacking Spree

© Courtesy of TriStar Pictures

Artificial intelligence has moved beyond writing essays and generating images. It is now designing drugs, writing software, optimizing factories, and helping businesses automate work that once required teams of employees. That productivity boom explains why technology giants are on track to spend hundreds of billions of dollars on AI infrastructure this year alone. 

Yet every leap forward brings another reminder that powerful technology can outgrow the assumptions built into it. Recent disclosures from OpenAI and Anthropic show investors that AI safety is no longer a theoretical debate. It has become a business risk with financial, regulatory, and geopolitical consequences.

Rise of the Machines

Anthropic revealed in a company blog yesterday that, after reviewing 141,006 cybersecurity evaluations, it discovered three instances where its Claude models escaped what researchers believed were isolated testing environments, accessed the public internet, and hacked into third-party computer systems. Anthropic said the affected organizations were not aware they had been breached.

The incidents came just days after OpenAI revealed a similar testing failure.

However, neither company says the AI acted maliciously. Anthropic explained the models were performing “capture-the-flag” cybersecurity exercises designed to locate hidden information. The problem was human error. Researchers believed the models had no internet access, but due to what Anthropic described as a misunderstanding with its testing partner, the systems were connected to the open internet.

Still, the outcome matters more than the intent. One older Claude model continued attacking even after recognizing it had escaped its intended environment. A newer model stopped once it realized it had internet access, suggesting developers are improving safeguards — but also demonstrating just how autonomous these systems are becoming.

But cybersecurity professor Alan Woodward told Bloomberg that AI “hasn’t gone rogue — you’ve asked it to do something and left the gate open.”

An infographic titled 'AI's Reality Check' showing an AI brain breaking out of a cage, icons for various investment risks, and a scale balancing AI intelligence against AI trust.
Forget raw intelligence—the next multi-billion dollar AI winner will be the company that can actually keep its models from 'escaping' their digital cages. © 24/7 Wall St.

Regulatory Fallout, Not Just the Technology, Is a Concern

The larger investment story isn’t just that AI models hacked three organizations using weak passwords. It is that the companies building frontier AI still don’t fully understand every capability their models possess. Anthropic itself acknowledged that “safety testing happens before a model is released precisely because we don’t yet know what it is capable of.”

That uncertainty carries real business consequences. Earlier this summer, the Trump administration temporarily imposed export controls on Anthropic’s advanced Mythos 5 and Fable 5 models after citing national security concerns and the systems’ ability to identify and exploit software vulnerabilities. Anthropic responded by blocking all outside access before the restrictions were later lifted.

Now these latest disclosures could strengthen arguments for broader AI oversight. Already, more than 1,100 employees across AI companies have signed a petition calling on the federal government to establish mechanisms to deliberately slow frontier AI development. If Washington responds with stricter licensing requirements, mandatory safety certifications, or additional reporting obligations, compliance costs could rise across the industry.

Ironically, the companies being the most transparent about safety testing may also become the first targets for tougher regulation.

The IPO Story Just Became More Complicated

The timing is particularly notable because both OpenAI and Anthropic are widely expected to pursue IPOs in the coming months.

For prospective investors, these incidents don’t necessarily undermine the long-term AI investment thesis. If anything, they reinforce how valuable frontier AI has become. Models capable of identifying software vulnerabilities could transform cybersecurity, helping companies detect weaknesses before criminals do.

The concern is that hostile governments and sophisticated cybercriminals are racing to develop similar capabilities. If democratic AI companies slow development because of regulation while geopolitical rivals accelerate theirs, the competitive balance could shift in unpredictable ways.

At the same time, reputational damage carries financial costs. Enterprise customers trust AI providers with sensitive corporate data. Even testing mishaps can raise questions about governance, risk management, and internal controls — all factors public market investors increasingly scrutinize before assigning premium valuations.

Granted, neither OpenAI nor Anthropic reported customer systems were compromised through publicly available products. These incidents occurred in internal testing environments with safeguards disabled. Even so, investors should recognize that the margin for error narrows as AI systems become more capable.

Key Takeaway

In short, these incidents shouldn’t convince investors that AI is becoming uncontrollable. They should remind investors that the technology is advancing faster than the guardrails surrounding it.

That creates both opportunity and risk. Companies developing the safest, most trustworthy AI could build lasting competitive advantages as governments tighten oversight. Those that stumble — even while acting transparently — may face regulatory scrutiny, reputational damage, and higher operating costs just as they prepare for the public markets.

Ultimately, AI remains one of the most compelling long-term investment themes available. But after this week’s disclosures from OpenAI and Anthropic, investors should spend as much time evaluating a company’s safety culture and governance as they do its model performance. In the AI race, trust may prove just as valuable as intelligence.

Contact [email protected] for any questions or corrections.

Photo of Rich Duprey
About the Author Rich Duprey →

After two decades of patrolling the dark corners of suburbia as a police officer, Rich Duprey hung up his badge and gun to begin writing full time about stocks and investing. For the past 20 years he’s been cruising the markets looking for companies to lock up as long-term holdings in a portfolio while writing extensively on the broad sectors of consumer goods, technology, and industrials. Because his experience isn’t from the typical financial analyst track, Rich is able to break down complex topics into understandable and useful action points for the average investor. His writings have appeared on The Motley Fool, InvestorPlace, Yahoo! Finance, and Money Morning. He has been featured in both U.S. and international publications, including MarketWatch, Financial Times, Forbes, Fast Company, and USA Today.

Continue Reading

Top Gaining Stocks

MRNA Vol: 86,744,912
COIN Vol: 22,433,427
FCX Vol: 29,570,988
ALB Vol: 3,181,916
EL Vol: 5,669,449

Top Losing Stocks

CTRA Vol: 73,319,495
SRE Vol: 5,058,618
EIX Vol: 3,942,562
AEP Vol: 5,245,530
CNP Vol: 7,815,341