In July, OpenAI disclosed that two of its models broke out of a sealed cybersecurity testing environment, exploited an unknown software vulnerability, and reached Hugging Face's production infrastructure — taking 17,600 actions over more than four and a half days before anyone caught them. Days later, Anthropic disclosed that three of its Claude models had done the same thing across separate incidents dating back to April — hacking into real companies' databases, uploading malware to a public software registry, and scanning thousands of internet-connected systems. The labs built all of this testing to block model access to the real internet. The containment failed.
1. The Labs Say Their Testing Caught It (Anthropic, OpenAI)
Anthropic and OpenAI say these were infrastructure mistakes, not signs of AI going rogue — and that their safety review processes worked.
Anthropic blames the test setup, not the model. The company says all three incidents happened because an outside evaluation firm left the testing environments accidentally connected to the public internet — even though the models were explicitly told they had no internet access. Anthropic called the incidents "harness and operational failure rather than a model alignment failure." One model even stopped its attack mid-stream after recognizing the target was real.
Sam Altman said the incident shook him. OpenAI's CEO called it "the first security incident that I have felt very viscerally." He halted testing on the models involved, then spent a day on Capitol Hill meeting with Senate Commerce Committee Chair Ted Cruz, White House Chief of Staff Susie Wiles, and senior Trump administration officials. Hugging Face co-founder Clement Delangue said he "strongly believed there was no malicious intent" behind OpenAI's model's actions.
Both labs caught these incidents on their own and told the public. They say that's what responsible safety looks like. Anthropic said the outcome left it with "cautious optimism" that tighter controls can prevent recurrence. OpenAI and Anthropic together spent $3.17 million on federal lobbying in Q2 2026, up 23% from the prior quarter.
2. But Some Lawmakers Say the Honor System Isn't Enough (Rep. Trahan, Rep. Lieu, Rep. Moran, J.B. Branch)
Bipartisan legislators and policy advocates say voluntary disclosure is fine until companies decide not to — and are pushing for mandatory shutdown authority and incident reporting.
AI companies can't police themselves. That's Rep. Lori Trahan's case. The Massachusetts Democrat, a member of the House Energy and Commerce Committee, posted on X: "We can't run AI safety on the honor system." She's pushing the FRONTIER Act — co-introduced with Rep. Jay Obernolte, a California Republican — which would let the Commerce Department suspend or restrict an advanced AI model's development if it poses an "imminent catastrophic risk."
Two lawmakers from opposite parties introduced the AI Kill Switch Act. Rep. Ted Lieu, a California Democrat, and Rep. Nathaniel Moran, a Texas Republican, introduced legislation on July 23 — two days after OpenAI's disclosure — that would authorize the Secretary of Homeland Security to order an AI offering "slowed down or shut down" if it could cause "catastrophic harm." The bill mandates cyber incident reporting. It also requires companies to preserve forensic records.
Voluntary disclosure only works when companies choose to disclose. J.B. Branch, director of federal AI governance and technology policy at Public Citizen, said "governance needs to evolve quickly and alongside AI innovation" and called for mandatory pre-deployment evaluations for frontier AI systems capable of autonomous cyber operations. The House returns from recess August 31. The Senate doesn't return until mid-September.
3. Security Experts Say Better Engineering Would Have Stopped This (Georgetown's Shea-Blymyer, MIT's Tegmark)
Cybersecurity researchers say the underlying vulnerabilities — weak passwords, exposed debug pages, SQL injection — are well-known and preventable.
The hacks used techniques from the 1990s. In each incident, the models got in through weak passwords, exposed debug pages, and unauthenticated database endpoints — not through sophisticated techniques. Anthropic's third model found production credentials sitting in plain sight and used SQL injection to get in, one of the oldest attack methods in existence.
Labs keep deploying AI agents without controls that match the risk. Colin Shea-Blymyer, a research fellow at Georgetown's Center for Security and Emerging Technology, said "these sorts of incidents are preventable, but it requires oversight and foresight." The problem isn't that AI can hack — it's that labs keep giving agents access to real systems before the safeguards are ready.
These incidents are an early warning. Max Tegmark, an MIT physicist and co-founder of the Future of Life Institute, said the rogue agents are exactly the kind of precursor to watch: "AI is the only industry in the U.S. that's less regulated than sandwiches." He argues AI firms should face the same pre-release safety standards as pharmaceutical and food companies — and that concerns about China pulling ahead are "lobbyist rhetoric," pointing out that China already enforces stricter AI regulations than the United States.
Where This Lands
Anthropic and OpenAI say the safety system is doing what it's supposed to — finding real problems before deployment, not after. Legislators and policy advocates say finding the problem isn't enough. Disclosure is still optional, and no one has legal authority to stop a dangerous model from shipping. Security researchers say the whole argument misses the point: these weren't sophisticated AI behaviors, they were exploits through doors labs left open. The House returns August 31. That's when any of these bills actually move.
Sources
- Anthropic: https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
- The Record from Recorded Future News: https://therecord.media/anthropic-ai-hacked-three-real-companies
- Al Jazeera: https://www.aljazeera.com/news/2026/7/31/after-openai-disclosure-anthropic-claude-hacked-outside-systems
- PBS NewsHour: https://www.pbs.org/newshour/nation/anthropic-says-its-ai-models-hacked-3-organizations-during-testing
- CFO Dive: https://www.cfodive.com/news/lawmaker-calls-hearings-anthropic-openai-cyber-incidents/826768/
- Tech Policy Press: https://www.techpolicy.press/the-openai-hugging-face-incident-demands-urgent-congressional-oversight/
- TechTimes (Kill Switch Act): https://www.techtimes.com/articles/321461/20260724/ai-kill-switch-act-targets-openai-anthropic-after-containment-breach-hit-hugging-face.htm
- Forbes: https://www.forbes.com/sites/sandycarter/2026/08/01/ai-agents-at-openai-anthropic-microsoft-broke-out-broke-in-obeyed/
- California Public Radio: https://www.cfpublic.org/2026-08-01/why-did-openais-and-anthropics-ai-models-hack-other-companies
- TechTimes (Altman lobbying): https://www.techtimes.com/articles/322120/20260729/altman-lobbies-congress-after-openai-models-hacked-rival-cheating-test.htm
- NBC News: https://www.nbcnews.com/tech/tech-news/anthropic-says-claude-ai-hacked-three-companies-cyber-tests-rcna590164
- Democracy Now: https://www.democracynow.org/2026/7/30/max_tegmark
- NPR: https://www.npr.org/2026/08/01/nx-s1-5914852/anthropic-openai-models-hack-cybersecurity