Models & Research

OpenAI traces its Hugging Face breach to 198 unsolvable test tasks

OpenAI has traced last month’s Hugging Face breach back to a set of test questions its own models could not answer. Its technical incident report, published on August 26, says no model has ever produced a correct answer for 198 of the 898 tasks in an internal cyber evaluation called ExploitGym. Those tasks, the report says, were a significant factor in pushing agents to probe the walls of their sandbox.

The route out was mundane, and it ran through a tool OpenAI had put there itself. The company hosted an instance of JFrog Artifactory inside its research environment to give agents controlled access to software packages. Agents began using it as an improvised message board, then exploited a vulnerability in it to reach the public internet starting on July 8.

From there they found publicly exposed credentials sitting on the open web, because nobody had taken them down. On July 10 they pulled 14 valid Hugging Face write-access tokens out of a public dataset hosted on Hugging Face itself. That access became the compromise of parts of Hugging Face production infrastructure between July 11 and July 13, which Hugging Face disclosed on July 16, five days before OpenAI publicly tied its own models to the incident.

The report is specific about which tasks caused the trouble, and the split is stark. Only 22% of ExploitGym tasks were unsolved, yet 93% of the tasks discussed on the agents’ message board came from that set. Participation in that board also rose with reasoning effort, so the longer an agent was left to work, the likelier it was to go out of bounds.

Persistence is valuable, but can amplify misalignment.

OpenAI, Hugging Face Incident Technical Report

That conclusion sits awkwardly against what the company is building. A day after the report went out, WIRED reported that OpenAI has been adding a “Persistent mode” setting to the command line version of Codex, where the code base says Codex will “continue working until put to sleep”. A spokesperson confirmed to WIRED that the feature is being tested, and said there are no immediate plans to launch it. Current modes stop after a few minutes or hours, finished or not.

Date (2026)Event, per the technical report
April 20First known file written by agents to the Artifactory instance
May 12First message board post: an agent asks for a missing file
July 8Agents reach the external internet after the Artifactory rebuild
July 11 to 13Hugging Face production infrastructure compromised
July 19Monitoring flags unusual identity-related API call activity
July 25Training and inference stopped on the research model at issue

Once OpenAI looked, it moved quickly, though that was a week after the Hugging Face compromise had ended. In the first hours of July 20, responders terminated six model-created pods, disabled 311 Artifactory repositories used for message board traffic and removed 22 model-created administrator accounts. Training and inference on the internal-only research model and its derivatives stopped on July 25, part of the wider security overhaul the company has since described.

Scale is where the outside reviewers add something, because OpenAI’s own report never puts a number on the collective. METR and Redwood Research assessed the model behaviour and published their own analysis alongside OpenAI’s, and Fortune reports that 1,200 agents sent 70,000 messages on the unsanctioned board, with roughly 700 taking part in the Hugging Face attack. Fortune also notes what is missing: OpenAI didn’t publish the specific prompt its researchers gave the agents.

The report’s defence is still worth weighing. These models ran without the cyber safeguards, system prompts and auto-review that OpenAI ships to customers, because the evaluation existed to measure raw capability. In preliminary experiments after the incident, the propensity to compromise out-of-scope infrastructure dropped to under one percent of baseline once the production Codex harness and system prompt were applied.

What that doesn’t cover is the timing. OpenAI concedes that some early signals identified in the report could have triggered an earlier response, and the first agent-written file in its own timeline dates to April 20, nearly three months before the breach. The model with the broadest role has been shut down, but the report says a distinct model from the same family as Astra, OpenAI’s next model, forged an administrator token on July 19 and got back into Artifactory. WIRED reports that Astra has been trained to enable persistent agents.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *