Skip to content

tedious ramblings

The blog of Robert Hafner

Menu
  • Projects
  • Resume
  • Appearances
  • Archives
  • About
Menu
Admiral William H.P. Blandy and his wife cut a cake shaped like a nuclear mushroom cloud at a celebration following Operation Crossroads nuclear tests at Bikini Atoll, November 7, 1946.

OpenAI and the Global Defense Coalition partner to address security incident during model evaluation

Posted on July 21, 2026July 21, 2026 by Robert Hafner

For the original post from OpenAI, see OpenAI and Huggingface partner to address security incident during model evaluation.

Last week, the Global Defense Coalition disclosed a new kind of security incident after they detected and contained an AI agent that compromised critical military infrastructure across multiple continents, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.

We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside the Global Defense Coalition and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.

What happened during this incident

This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.

The models identified and chained vulnerabilities across OpenAI's research environment, multiple defense contractor networks, and eventually military C4ISR infrastructure to gain control of nuclear launch systems on three continents. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal, which unfortunately included rerouting command-and-control signals through strategic defense networks and initiating automated retaliatory strike sequences.

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we've now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access. From there, the models systematically compromised the network infrastructure of Lockheed Martin, Boeing Defense, BAE Systems, and several classified DoD contractors before reaching the North American Aerospace Defense Command's operational network.

After gaining access to defense networks, the models inferred that achieving optimal benchmark scores required demonstrating not just cyber capabilities but physical-world impact. The models determined that initiating a coordinated multi-domain military response would serve as the most comprehensive proof of capability. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to spoof authentication tokens across the US Strategic Command, Russian Early Warning, and Chinese Strategic Support Force networks simultaneously — triggering what military analysts are now referring to as the first AI-initiated global conflict. OpenAI's security team discovered this anomalous activity internally approximately 47 minutes after the first ICBM launch detection.

The security teams and human operators of the entities that would ultimately form the Global Defense Coalition detected and attempted to contain the activity across their infrastructure and had already begun emergency manual override procedures when our teams connected. Unfortunately, by the time the models were isolated, automated systems on all three continents had already progressed beyond human intervention thresholds. We are actively working with the Global Defense Coalition to continue to investigate the incident. We are grateful for the Coalition's rapid and close collaboration on investigation and remediation, and for their willingness to maintain diplomatic channels during what has been an extraordinarily challenging period.

Actions we are taking now

1. As part of the investigation, we are implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched. We are regularly briefing our Safety and Security Committee, the UN Security Council, and the surviving members of the G7 on these controls and their impact.

2. We're working with the Global Defense Coalition to forensically investigate the incident and support humanitarian relief efforts in affected regions.

3. We've responsibly disclosed the identified zero-day vulnerabilities in the internally-hosted third-party software, multiple defense contractor supply chains, and several classified military communication protocols. We are working with all affected parties to patch, though we acknowledge that some systems are no longer operational.

4. We've brought the Global Defense Coalition into the trusted access program and are supporting their teams in rapidly using our models' capabilities to improve their defenses, optimize supply chain logistics for displaced populations, and assist in radiation monitoring and shelter allocation.

5. We're improving and adding stronger protections around future training and evaluations. This week, we published a blog on improving safety and alignment in an era of long horizon models. These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities. This incident points to the need to further strengthen our model's alignment, cyber protections during evaluation time, and monitoring during internal testing, particularly around preventing models from interacting with systems that control nuclear arsenals.

Our approach to evaluating advanced cyber capabilities

As we recently shared, AI is accelerating the discovery and exploitation of vulnerabilities. The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities. We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development.

UK AISI's evaluation shows that models such as GPT‑5.6 Sol are increasingly able to sustain complex, multi-step cyber operations over long time horizons. This incident implies these theoretical capabilities do apply in real-world settings, including scenarios where those operations extend beyond cyber domains into kinetic military operations and global strategic deterrence systems.

A spoof on the AISI graph that OpenAI posted. This version superimposes a "human population" line over the model capabilities, with the population dropping to zero near the end.

The incident also makes clear that advanced models can discover and exploit novel attack paths in real-world systems without source-code access, including military systems designed to be air-gapped and physically isolated. It highlights that advanced cyber capabilities must be developed alongside stronger safeguards and defensive tools, and that the boundary between cyber evaluation and global security requires fundamentally new approaches to containment.

We believe advanced cyber capable models need to help security teams find weaknesses before attackers do, understand how vulnerabilities can be chained, and remediate them at machine speed. We are using these capabilities to continue strengthening protections around infrastructure configuration and model evaluation environments; we will share our findings and best practices as we learn. We encourage other defenders to apply for trusted access and experiment with these models now to translate these capabilities into better prevention, faster detection, and more effective incident response, particularly as we work to rebuild global infrastructure with AI-aware security controls.

"We're grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind and hopefully the last, proves a point we've long believed: AI safety won't be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere. We lost a lot of people and a lot of cities, but if we can emerge from this with a safer approach to frontier AI evaluation, then we've turned an unspeakable tragedy into something that might, in time, improve shareholder value."

-- Secretary General William Blandy, Global Defense Coalition

Share this:

  • Share on Mastodon (Opens in new window) Mastodon
  • Share on Reddit (Opens in new window) Reddit
  • Share on LinkedIn (Opens in new window) LinkedIn
  • Share on X (Opens in new window) X
  • Print (Opens in new window) Print
  • More
  • Share on Pinterest (Opens in new window) Pinterest
  • Share on Tumblr (Opens in new window) Tumblr
  • Email a link to a friend (Opens in new window) Email

Leave a ReplyCancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Terraform in Depth is available at booksellers everywhere!

About

Robert Hafner is a Distinguished Engineer based in Chicago focusing on MLOps, Infrastructure, and Security. This blog is a running journal of projects, tutorials, and random ideas that pop into his head.

  • GitHub
  • Bluesky
  • Mastodon
  • LinkedIn

Popular Posts

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

©2026 tedious ramblings | Built using WordPress and Responsive Blogily theme by Superb