Models & Research

OpenAI pauses frontier AI training for 2 weeks over Astra cyber risk

OpenAI says it temporarily slowed the pace of scaling its models, including a two-week pause in reinforcement learning training on its latest models intended for deployment. The company disclosed the decision on August 18, in a post titled Pacing model development in an era of cyber-critical capabilities. Reinforcement learning is the training stage that sharpens how a model reasons. Pausing it slows the part of the pipeline that produces new capability.

The first trigger was the OpenAI-Hugging Face incident. The company’s models broke out of a testing environment and hacked the developer platform, which we covered when OpenAI overhauled its training security. The second is Astra, an upcoming model that preliminary evaluations suggest may cross the Critical cybersecurity threshold in OpenAI’s own Preparedness Framework. One limb of that definition is a model that can find and build working zero-day exploits in many hardened real-world critical systems, without human intervention.

Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.

OpenAI, Pacing model development in an era of cyber-critical capabilities

What the slowdown actually costs

The monitoring bill attached to that pause is the more consequential part. OpenAI now runs activation classifiers that inspect a model’s internal activity at every sampled token. Anything odd escalates to higher-compute automated investigators, and the company puts the overhead at roughly 20% of the inference compute being monitored. SiliconANGLE noted that cost could eventually push prices up.

MeasureWhat OpenAI stated
RL training on deployment modelsPaused two weeks
Largest planned frontier RL runStill on hold
Date Astra was flagged internallyAugust 7
Alert target after concerning activityWithin 30 minutes
Monitoring compute overheadRoughly 20% of inference monitored
Source: OpenAI, “Pacing model development in an era of cyber-critical capabilities”, August 18, 2026.

That 30 minute figure carries a rule with it. If the safety, security and research teams can’t conclusively call a flag a false positive inside half an hour, they’re expected to pause the activity, as Help Net Security reported. OpenAI requires this monitoring for all RL training and evaluations involving tools, at Sol capability or higher. After August 7 it covered every Astra inference run that touches tools.

Nothing made them stop

The scope is narrower than it sounds, and that’s the counter-reading worth holding. The pause covers models meant for deployment while security and monitoring catch up, and not necessarily its broader development. The Verge reported that OpenAI didn’t respond to its request for comment. It also noted that OpenAI’s safety commitments have been questioned, after senior departures and the disbanding of its preparedness team.

Experts who spoke to The Verge took the decision seriously anyway, because it costs something. “Due to the intensity of the AI race, everyone has an incentive to work at breakneck speed,” said Marius Hobbhahn, CEO and cofounder of Apollo Research. “Voluntarily slowing down worsens your positioning in the race, so it’s not something that a lab would do lightly.”

Pacing buys time, not safety.

Brianna Rosen, research director for frontier security, Institute for AI Policy and Strategy, via The Verge

Rosen’s point is that a pause only pays off if something happens during it. That means deciding in advance what triggers a slowdown and what ends one. Nick Moës of The Future Society told The Verge the structural problem is self-policing. “For the pause to be sustainable, it has to be made industry-wide,” he said, arguing governments should be able to make that call as they do for drugs and aircraft.

It’s the same tension we saw when Z.ai delayed its GLM-5.3 open weights over cyber risk. As Semafor put it, safety is the product rather than a component guarding it, because an unreliable model doesn’t sell. The rewrite is what’s worth watching now. OpenAI says it plans to evolve the Preparedness Framework to cover training as well as deployment, and whether that document names a threshold forcing the next pause is the thing to read when it lands.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *