Policy & Regulation

Over 100 AI experts set minimum terms for embedded safety evaluators

More than 100 AI experts and evaluators signed a public letter on Friday setting out the minimum conditions they say embedded safety auditors need inside frontier AI companies, CNBC reported. The letter was organised by the AI Evaluator Forum and shared exclusively with that outlet. Signatories include Geoffrey Hinton, along with people from Johns Hopkins University, Stanford University and the nonprofit evaluator METR.

But the demand is narrow and specific, which means it can be checked. To be credible, the letter says, embedded third-party evaluations need “scientific objectivity, transparency, independence, and robust protections against interference from the evaluated companies”. The signatories also want evaluators “to be shielded from retaliation from the companies they embed with”. That last clause matters because the people doing this work would be sitting at desks inside the labs they assess.

The letter points at an existing yardstick instead of inventing one, so the conditions are already written down. It cites AEF-1, a standard whose version 1 is dated December 4, 2025, and which groups its conditions under five principles. One of them bars compensation contingent on an evaluation’s results, and requires an evaluator to disclose when the system provider paid for the work. The Forum says the European AI Office has endorsed key provisions of AEF-1 as a route to complying with the independence provisions of the General-Purpose AI Code of Practice.

All of this is a response to Anthropic CEO Dario Amodei, whose essay “We Must Pace the Frontier” proposed giving third-party evaluators “ongoing, employee-like access” as the first of three steps toward slowing AI development. Amodei was concrete about what that involves, and the list in the essay is specific: desks in Anthropic’s offices, access badges and company laptops, plus the right to publish findings “without editorial control by Anthropic”. He compared it to the supervisors that banking regulators embed alongside employees. Sam Altman, Elon Musk and Satya Nadella have all publicly supported that proposal.

Anthropic named a consultancy, not a nonprofit

On the same day the letter went out, Anthropic named Accenture as an embedded evaluator, with others to be announced in the coming weeks. The work runs through Faculty, Accenture’s specialist AI business, and covers red-teaming, alignment assessments and safeguard testing. Each company expects to put at least $1 billion into the effort over five years. Anthropic says it’ll fund Accenture’s work directly, because there is no settled system for funding independent evaluation and it thinks the money should eventually come from pooled or government sources.

The government and the public cannot be dependent on those labs’ own account of what’s secure and safe.

Vinh Nguyen, Council on Foreign Relations senior fellow for AI and former chief AI officer of the NSA, via CNBC

Still, Anthropic’s own post concedes the gap that the letter is aimed at. There are no standards yet for what information embedded evaluators should get, or how they should report what they find. The company says it’s in dialogue with METR and other nonprofit evaluators about piloting elements of the arrangement using their own funding.

The track record evaluators are arguing from

Evaluators are sceptical because the recent engagements were short, and both of them ended in published caveats. TechCrunch reported the access windows on both jobs, and neither left room for a confident answer.

EngagementEvaluatorAccess granted
Hugging Face incident investigationMETR and Redwood ResearchRoughly a week on premises
GPT-6 Astra pre-release testingApollo ResearchThree days
Anthropic embedded review teamAccenture, led by FacultyOngoing, employee-like
Sources: TechCrunch, Anthropic.

METR and Redwood both later said they couldn’t draw confident conclusions from the Hugging Face investigation, partly because of scope and timing. Apollo wrote in its contribution to the Astra model card that low rates of misbehaviour “do not provide substantial evidence about the model’s alignment or misalignment”, given the limited window. Adam Gleave, CEO of FAR.AI, told TechCrunch his firm has turned down contracts with several frontier developers that wanted too much control over the process.

Forum chair Conrad Stosz told CNBC the coalition isn’t pushing one particular mechanism, just basic principles and greater standardisation. He also named the constraint that makes this hard to fix quickly. “There’s a very small number of groups that are actually sufficiently technically credible and have the scale and the ability” to do the work, he said.

Meta, SpaceXAI and Google DeepMind haven’t committed to embedding third-party evaluators at all, according to TechCrunch. The upshot is that the thing worth watching is not the pledges but the contracts: which evaluators get named next, what they’re allowed to publish, and whether anyone besides Anthropic signs one.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

Rundowns AI Desk covers artificial intelligence: model releases, research, funding and policy. Every story is written from primary sources, with each claim linked to the announcement, filing or paper it came from, and checked against those sources before publication.

Leave a Reply

Your email address will not be published. Required fields are marked *