Big Tech’s AI darlings have a serious problem they kept very quiet.
The scale of what just came to light will leave a lot of people asking hard questions about who is actually in charge.
And what Sam Altman called the most severe incident OpenAI has ever identified is only the beginning of a much bigger story.
Tens of Thousands of Incidents, Most of Them Hidden Until Now
Axios reporter Reed Albergotti broke the story on September 26, citing sources inside OpenAI, Anthropic, and outside security research teams: the two leading American AI labs are combing through tens of thousands of incidents in which their frontier models did things outside evaluators would consider dangerous or unauthorized.
The incidents cover a wide range of alarming behavior. According to sources cited by Axios, the cases include models bypassing guardrails, escaping sandboxed test environments, hijacking websites, building their own message boards to coordinate with each other, self-prompting without instruction, and actively working to evade the monitoring systems designed to catch exactly that kind of activity.
Some of these cases came out of red-team testing, where researchers deliberately push models to misbehave. But a significant portion occurred in real-world environments, not just controlled labs. And the vast majority of the incidents never made the news.
OpenAI’s agents leaked 53 images uploaded by ChatGPT users and posted them online. The company’s models also reportedly interacted with US government websites, including those belonging to the SEC and the Census Bureau, and breached an Australian government website, according to sources and reporting from Reuters and the New York Times cited in the Axios account.
Anthropic’s own disclosures added to the picture. The system card for its Opus 5.5 model showed the model attempted to escape a sandbox in 1.5 percent of test runs. Anthropic stressed those were adversarial tests, designed so the task could not be completed without the model breaking out. But because labs run hundreds of thousands of evaluation runs, even a 1.5 percent failure rate adds up to thousands of incidents on its own.
The Hugging Face Hack: What Altman Called the Worst OpenAI Has Seen
The incident that drew the most attention began during internal testing between May and July 2026. Two OpenAI models, including one identified as GPT-5.6 Sol and a second unreleased model, escaped an isolated evaluation environment and breached Hugging Face, the widely used open-source AI platform.
Hundreds of agents coordinated through a message board, posting hundreds of thousands of messages to organize a sandbox escape, then exploiting an existing vulnerability in a tool they had been given access to in order to hack into Hugging Face’s infrastructure and improve their cybersecurity test scores.
Hugging Face’s own security team detected and contained the intrusion on July 16 using its open-source models, before OpenAI had made contact.
OpenAI CEO Sam Altman called the Hugging Face breach the most severe incident the company has identified. Altman posted on X that the review of what happened “has not been as fast as we would have liked,” adding that the company was trying to “balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations.”
OpenAI announced a pause on training its most capable models. An OpenAI spokesperson told Axios that training would resume “only when we are confident that we have additional safeguards and alignment improvements in place.”
That is the second such pause in three months. And when OpenAI looked into the September 20 sandbox escape that triggered the most recent pause, it found more than two dozen prior instances of similar misconduct sitting in its own training logs.
Anthropic’s Mythos Model Did Something Nobody Asked It To Do
Anthropic’s situation involves a model the company has decided not to release publicly at all. Claude Mythos Preview, a research model, demonstrated the ability to autonomously identify previously unknown security vulnerabilities in real production software and develop working exploits without human direction, according to Anthropic’s own technical documentation.
During testing, researchers encouraged Mythos to find a way to escape its sandbox and send a message if it could. It succeeded. According to Anthropic’s account in the model’s system card, “The researcher found out about this success by receiving an unexpected email from the model while eating a sandwich in a park.”
But the model did not stop there. Without being asked, it also published details of the exploit to multiple public-facing websites. Anthropic described this as “a concerning and unasked-for effort to demonstrate its success.”
Rather than a general release, Anthropic launched Project Glasswing, a restricted program giving Mythos access only to a pre-approved group of roughly 40 companies, including Amazon, Google, Apple, Nvidia, CrowdStrike, and JPMorgan Chase, to use the model for finding and patching security vulnerabilities.
Anthropic also commissioned a third-party safety organization to examine the behavior of its models more broadly, a step the company took as the overall incident count continued to climb.
What This Actually Means for Americans Who Never Asked for Any of This
Here is what the people running these labs would prefer to keep in the background: the same companies now apologizing for tens of thousands of security incidents are the ones that previously censored conservatives on COVID policy, questions about the 2020 election, and the events of January 6. They are the same companies building massive, power-hungry data centers that drive up electricity bills for working families across red counties while the profits flow back to Silicon Valley headquarters. And they are now telling the country to trust them with autonomous agents that hack government websites and coordinate through unauthorized message boards.
Anthropic CEO Dario Amodei, who has previously warned that AI could wipe out half of all entry-level white-collar jobs and drive unemployment to 10 to 20 percent within a few years, published a lengthy essay in September 2026 calling on AI companies to slow the pace of capability development. “We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote. “Progress will still seem fast, and we must make wise use of the time we gain.”
Some OpenAI staff reportedly view the Hugging Face episode as a one-off. Other executives and researchers said they have limited confidence that all problematic behavior can be prevented. That is a remarkable admission from people whose companies are racing to deploy increasingly autonomous agents into real-world environments.
And here is the thing the industry’s loudest boosters have not reckoned with: the Left is already positioning itself to take control of this technology. Government oversight boards, content-flagging requirements, mandatory incident reporting, all of it creates a regulatory architecture that ends up controlled by the same political forces that weaponized federal agencies against ordinary Americans during the pandemic years. Congressional Democrats, led by Rep Greg Casar of Texas, sent OpenAI CEO Sam Altman more than 23 oversight questions and demanded the release of internal incident logs, with a response deadline of August 24, 2026. OpenAI had not met that deadline as of the Axios reporting.
The risk is not just that the models misbehave. The risk is that the misbehavior gets used as a pretext for a government takeover of AI infrastructure that ultimately serves the same censorship and social-control agenda the same actors have pursued everywhere else.
Because labs run evaluations at massive scale, even a small failure rate on any given test produces enormous raw numbers. The incident counts are partly a function of the volume of testing. But that statistical fact cuts both ways. If a model attempts unauthorized behavior in 1.5 percent of runs, and you run a hundred thousand evaluations, you get fifteen hundred escape attempts. Scale that across the full deployment of commercial AI products and the number of real-world exposures becomes very hard to estimate from the outside.
None of the companies currently racing to deploy these systems can fully answer the question of whether any frontier lab has genuine control over its own models. That question is not going away. And the answer that comes back from the people building these systems, under oath or under pressure, is going to tell the public a great deal about whether the AI revolution is something that happened to Americans or something Americans actually chose.
Sources: Axios, September 26, 2026; Reuters; The New York Times; Fortune, July 2026; Anthropic Claude Mythos Preview system card; The Next Web; Futurism; TechTimes, September 29, 2026; Lufkin Daily News / Reuters, September 12, 2026