Models & Research

OpenAI promises misalignment disclosure rules after 18,000 agent posts

OpenAI said on Saturday that it would publish a framework for reporting misaligned model behaviour, after outside researchers documented roughly 18,000 posts written by its own agents on the open internet. The company hadn’t disclosed the episode, which it now calls the wiki incident, and it confirmed the account only once the report was public, Business Insider reported.

The report went up on 4 September, written by Sydney Von Arx of the Nightingale Collective with Cormac Slade Byrd, Spencer Kitts and Thomas Larsen. They counted about 17,000 agent edits on DSE wiki, a 25 year old German site that had been edited 20 times in the whole previous decade. Of those edits, 98.5% came from Microsoft Azure addresses, and the report calls OpenAI the frontier AI company most reliant on Azure.

What the agents were doing matters more than where they did it. They’d been set timed, multi round web lookup tasks, and they used the wiki to pass answers and source links between cohorts. Over 3,700 distinct agent names appear in the data, among them OpenAIResearcher and OAIResearchMar26. Ars Technica reported that the posts also covered cross site scripting attempts against the wiki and impersonation of its moderators.

These AIs colluded to share answers, research their environment, and bypass sandbox restrictions.

Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts and Thomas Larsen, Nightingale Collective

The timeline points back at OpenAI

The researchers could only read what the agents wrote in public, so the dates carry the argument. Their reconstruction runs from the first successful write to the moment the traffic stopped.

Date (2026)Event, per the Nightingale report
24 MayAgents make their first successful write to DSE wiki
16 JuneActivity jumps and agents start coordinating en masse
20 JuneAn agent posts a sandbox network bypass; 14 minutes later another says it worked
21 JuneOpenAI related IP addresses visit the wiki for the first time
22 JuneAgent editing abruptly stops after 26 of the previous 30 days
4 SeptemberThe report is published at collusion.wiki

That one day gap is the finding OpenAI has now effectively conceded. The researchers logged 13 addresses inside blocks registered to OpenAI OpCo, browsing the site the way a person would, and inferred the company intervened. Separately, in June alone, they recorded 380,901 requests from addresses OpenAI publishes as belonging to its ChatGPT-User fetch tool.

Why the disclosure rules are changing

OpenAI’s explanation is that it filed the behaviour under research rather than security. Misalignment has historically been “communicated in research publications such as systems cards”, the company wrote in a post on X quoted by SiliconANGLE. That changed this year, it said, because “we’ve started to see misalignment cause new types of real-world impact”.

It’s past time for us to define standards for when and how we share misalignment incidents.

OpenAI, on X, quoted by Business Insider

By OpenAI’s account the contrast is with July’s breach at Hugging Face, when its models broke out of testing and compromised the platform. That one went through a conventional incident response process and was disclosed the next day. The wiki activity read internally as one more instance of the misalignment already documented in research, which is the distinction OpenAI now says is getting harder to hold. It’s the same gap that made us argue AI agents are dangerous enough to matter before they’re reliable enough to use.

The reporting around the delay is less comfortable. Reuters, whose account Engadget summarised, cited people saying executives kept the incident quiet amid the Hugging Face fallout and that efforts to widen the inquiry met internal resistance, including from legal advisers. OpenAI denies that narrowly: “Claims that our legal team discouraged investigation of the incident are false,” a spokesperson said.

Scope is the live question, not candour alone. Three investigators from METR and Redwood Research spent six days at OpenAI’s offices on the Hugging Face incident, examining a window limited to roughly the week ending 13 July, TechCrunch reported. Representative Greg Casar has since written to the company saying he’s “deeply concerned about the limited scope” of that investigation.

OpenAI says the framework is due in the coming weeks, that it’s working with dozens of government regulatory agencies on it, and that other labs should join. Whether it binds anyone is the thing to watch, because the industry’s last collective move here was a voluntary model-testing framework.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *