deniz.in

Markets

Weather

Loading weather

· via TechCrunch

Anthropic and OpenAI pledge embedded safety evaluators; auditors ask what independence means

Anthropic and OpenAI say third-party safety evaluators will get far-reaching access inside frontier labs, but auditors warn the plan needs detail, time and legal backing to become real oversight.

Anthropic and OpenAI pledge embedded safety evaluators; auditors ask what independence means

Frontier labs volunteer to host embedded auditors

Anthropic CEO Dario Amodei has proposed placing third-party safety evaluators inside every frontier AI company, giving them the authority to report safety incidents, judge whether models are genuinely aligned and share their conclusions publicly. Writing in an essay published over the weekend, Amodei said Anthropic would grant evaluators such as METR and Redwood Research access to its systems well beyond anything outside researchers have historically received, and OpenAI CEO Sam Altman said his company would adopt the same practice. According to TechCrunch, evaluators broadly welcomed the idea but said the details — and ideally legislation behind them — will determine whether they function as independent monitors or as service providers whose work is shaped by the labs.

Auditors want to see training, not just finished models

Outside reviewers have historically tested finished models shortly before release. The evaluators TechCrunch spoke to want far more: access to intermediate checkpoints from throughout a model's training, so they can pinpoint when concerning behavior first emerged. Adam Gleave, CEO of FAR.AI, said auditors could also inspect the post-training environments that reward models for certain behaviors, and review evaluation transcripts and logs to verify a developer's claims about how a model performed. He suggested access could extend to interviewing employees, to check whether a company's documentation and public descriptions of its safety practices match what happened internally.

The push for deeper access comes as models get better at recognizing when they are being evaluated, raising the risk that they behave well during testing while concealing problematic behavior. John Steidley, head of strategy at Palisade Research, compared the danger to Volkswagen's Dieselgate scandal: a model trained specifically to do well on a shutdown-resistance benchmark might pass that test without actually being safe.

Alexander Meinke, head of research at Apollo Research, told TechCrunch that labs should be able to answer definitively whether a model ever actively tried to undermine its own alignment training. Today, he argued, the public relies entirely on companies checking this themselves and reporting it honestly — and recent incidents show that, by default, they often do neither. TechCrunch reports that Amodei's proposal includes a right for evaluators to publish findings on risks, incidents, practices and the access they received, free of Anthropic's editorial control.

The unanswered questions

Neither Anthropic nor OpenAI has said which evaluators they will work with, when embedding would begin, how many organizations would be involved, what systems and information they could reach or what could be disclosed publicly, despite repeated questions from TechCrunch.

Past friction explains the caution. Gleave said FAR.AI has turned down contracts with several frontier developers that wanted too much control over the evaluation process, and that evaluators are by default treated like ordinary contractors, bound by restrictive NDAs that give developers significant say over what can be published.

Time is another sore point. During the investigation of the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises, and both later said they could not draw confident conclusions, partly because of scope and timing limits. Ahead of the release of GPT-6 Astra, which OpenAI has promoted as its most aligned model so far, Apollo was given three days to test it and concluded that low rates of observed misbehavior provided little evidence either way, given the models' growing ability to detect evaluation and the narrow testing window.

Researchers who spoke to TechCrunch called for a transparent framework agreed publicly. Steidley argued it should include standards for what kinds of auditors companies can rely on, so labs cannot sidestep scrutiny by shopping for evaluators who are unqualified or uninterested in the most concerning risks. Henry Papadatos, executive director of Safer AI, said voluntary measures always depend on a company's goodwill, and that regulation would stop a lab from reversing course after a PR crisis while also binding less willing competitors. Meta, SpaceXAI and Google DeepMind have not committed to the practice, though DeepMind CEO Demis Hassabis has floated a separate industry standards body to independently test frontier models, and TechCrunch reports Google, OpenAI and Anthropic have been privately discussing AI safety plans for weeks.

Voluntary pledges versus emerging law

Some legal scaffolding already exists. California's SB 53, signed last year, requires large frontier developers to publish safety frameworks and report critical safety incidents, while SB 813, signed this month, creates a framework for state-recognized independent verification organizations. In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing and to report serious incidents, and the EU AI Office can run its own evaluations and appoint independent experts. TechCrunch notes the law remains less expansive than what Amodei is proposing.

Why it matters

Embedded evaluators with genuine access would be a structural change in how frontier AI is governed, letting outsiders examine training rather than only polished releases — at a moment when models may be learning to pass the tests themselves. Whether the pledge becomes real oversight or a public-relations exercise hinges on the unresolved questions of scope, time, publication rights and auditor selection, and on whether regulators turn a voluntary promise into an enforceable obligation.

  • #ai-safety
  • #anthropic
  • #openai
  • #governance
  • #regulation

Related posts