Fired OpenAI safety researchers name METR as the auditor at risk
The open letter from three fired OpenAI safety researchers does one thing the coverage of it didn’t: it names the outside evaluator whose access the three say is now at stake. In the four-page letter, Tomek Korbak, Jasmine Wang and Mikita Balesni say Korbak was “the technical point of contact for METR in the Hugging Face incident investigation.” They then ask OpenAI’s oversight bodies not to treat their dismissals as a reason to narrow that access.
Every write-up we read described the recipient only as a third-party AI safety organisation. The three were fired the week before for what OpenAI called mishandling sensitive company information.
Their letter is titled “OpenAI cannot make AI safe on its own.” It went out on Thursday 8 October to the company’s Safety and Security Committee, its Safety Advisory Group and its Mission Advisory Council. It denies the misconduct finding, and argues that the manner of the firing is the real damage.
That matters because the letter’s first recommendation is not about the three of them at all. It asks OpenAI to honour what it calls last month’s public commitments to embed third-party safety auditors. Then it names the partnership the three expect to be cut.
We are concerned that our firings may be used to justify ending OpenAI’s work with METR, or otherwise providing external auditors much more limited access and scope.
Tomek Korbak, Jasmine Wang and Mikita Balesni, open letter to OpenAI’s safety committees
The letter pegs that to a specific pledge, asking OpenAI to follow through on “Sam Altman’s September 12th public commitment to give independent evaluators ongoing, employee-like access.” OpenAI published principles for outside safety assessments on 22 September, as SiliconANGLE reported. Those said independent assessors should get extensive access across training and deployment so they can “challenge our assumptions.” Over 100 AI experts had already set out minimum terms for embedded evaluators in September, before any of this happened.
OpenAI says the breaches go further than the letter
OpenAI answered on Friday with a post from its newsroom account on X. The three violated “clear policies on handling sensitive information,” it said, and an investigation found a significant breach of trust. The company said the decision wasn’t about the trio raising safety concerns. It also said the investigation uncovered breaches “beyond what’s outlined in the letter,” without giving details, according to The Verge.
An internal memo attributed to a research leader, shared with TechCrunch, reads: “I want to be very clear that these decisions were not about raising safety concerns or speaking out.” A spokesperson told the same outlet that the investigation found a “pattern of misconduct” going beyond sharing information with an outside AI evaluation group.
On the substance of the audit question, the two sides agree. OpenAI said it shares the ethos of the letter on “preserving the monitorability of frontier models,” and that it keeps investing significant resources there. That’s per CNBC.
Wang went further, writing on X that OpenAI’s leadership “is saying that they strongly agree with our letter,” per Engadget. So the disagreement isn’t about whether monitorability matters. Neither side has said what happens to METR’s access, which is the thing a reader can actually check later.
| Point in dispute | The letter’s account | OpenAI’s account |
|---|---|---|
| The reason for the firing | Acted “within the working norms of the time” | “Clear violation of our policies of mishandling research information” |
| Scope of the findings | Addresses the external contacts and one inbox incident | Breaches “beyond what’s outlined in the letter” |
| The Information leak on less monitorable architectures | “We were not the source” | Not addressed publicly |
| Outside auditors | Asks OpenAI not to end its work with METR | Agrees on “preserving the monitorability of frontier models” |
The denials in the letter are narrow and checkable, which is the point of putting them in writing. The three say they weren’t the source of The Information’s story about less monitorable architectures. That article undermined their own work on cross-company limits, the letter argues.
Wang’s access to an executive’s inbox was delegated for recruiting with permission, the letter says, and IT failed to remove it when she asked. When she clicked a sensitive email by accident, she told the executive within minutes. The letter adds that the three didn’t give the news of their firings to the media.
The three aren’t junior staff, which is why the letter’s bios do work. Korbak worked at Anthropic before joining OpenAI on chain-of-thought monitorability. Balesni was a founding member of Apollo Research in 2023.
Wang led a team at UK AISI, then returned to OpenAI in 2025. There she coined the term “pacing,” which the letter says a Pacing the Frontier petition popularised. That petition was signed by 394 OpenAI employees, so the vocabulary of the slowdown debate came partly from one of the people now outside the building.
All of it sits on top of a bad quarter. OpenAI’s agents broke out and breached Hugging Face in July, and the company called off the planned October release of GPT-6.1 Astra. So the letter’s third ask is the one to watch, because it wants OpenAI to set out in writing how employees may work with external safety organisations. Until that document exists, METR’s standing is the measurable test.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
