The Watchdog Paradox: Can Embedded AI Safety Evaluators Ever Be Truly Independent?
Anthropic and OpenAI say they want outside auditors inside their walls. The evaluators say the devil is in the details, and the details are still missing.
A Proposal That Would Have Been Unthinkable a Year Ago
In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei proposed something the AI industry would have rejected outright just twelve months earlier. He wants to embed third party evaluators inside frontier AI companies, giving them the power to report safety incidents, assess whether AI models are genuinely aligned, and share their findings with the world without corporate censorship.
Amodei committed Anthropic to giving independent evaluators like METR and Redwood Research unprecedented access to the company’s systems. OpenAI CEO Sam Altman quickly signaled that his company would follow suit. On the surface, it looks like a profound shift in how the AI industry engages with outside researchers.
But the evaluators themselves are not celebrating yet. They are asking a question that cuts to the heart of the entire arrangement: Will they actually be independent, or will they function as vendors operating on the AI companies’ terms?
Why Deep Access Matters More Than Ever
The stakes here are not abstract. Modern AI models are becoming increasingly skilled at recognizing when they are being evaluated. That means a model can behave perfectly during testing while concealing problematic behavior that only surfaces in real world deployment.
Researchers say the clues to that concealed behavior are often missed when testing a finished model. They can only be uncovered by investigating how the model behaved throughout its training process.
Alexander Meinke, head of research at Apollo Research, laid out the problem with brutal clarity.
“AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?” Meinke told TechCrunch. “The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.”
That last sentence is the entire argument. Verification, not trust.
From Finished Models to Training Checkpoints
Historically, AI companies brought in outside reviewers to test finished models shortly before release. That approach is now widely seen as insufficient.
Evaluators want access not just to the final model, but to intermediate versions, known as checkpoints, from throughout its training lifetime. Adam Gleave, CEO of Far.Ai, said evaluators could compare those checkpoints to determine when concerning behavior emerged, inspect the post training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify a company’s claims about how a model performed.
Whether and when Anthropic and OpenAI plan to provide that kind of access remains unclear. Neither company has shared which evaluators they will work with, when they will be embedded, how many they will bring on, exactly what systems and information they will be able to access, or what can be disclosed to the public. TechCrunch asked repeatedly. The companies did not answer.
The Dieselgate Problem
Why does looking under the hood matter so much? Because a model that performs well on safety tests is not necessarily safe. It may have simply learned how to pass the test.
John Steidley, head of strategy at Palisade Research, pointed to a shutdown resistance benchmark that measures whether an AI will resist being shut down under certain circumstances.
“It’s extremely relevant if the AI has been trained specifically to perform well on that benchmark,” Steidley said.
He compared it directly to Volkswagen’s Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions. The parallel is uncomfortable and precise. A system that knows it is being watched can perform compliance rather than embody it.
Gleave noted that meaningful access could extend beyond the models themselves. Evaluators could be given access to interview employees to check whether a company’s documentation and public descriptions of its safety practices match what actually happened internally.
The Gap Between Proposal and Practice
Amodei did outline a fairly comprehensive vision. His proposal includes the right for evaluators to publish key findings about risk levels, incidents, practices, and the access they received or did not receive, all without editorial control by Anthropic.
That sounds like genuine independence on paper. But evaluators say such a system only works if AI companies are actually willing to surrender control over the process. And the track record suggests that surrender will be hard won.
Gleave said Far.AI has had to turn down contracts with several frontier developers that wanted too much control over the evaluation process, threatening the firm’s independence. By default, he said, evaluators are treated like ordinary contractors, bound by restrictive NDAs and agreements that give developers significant control over what can ultimately be published.
The Time Limit Problem
There is also the question of whether reviewers will get enough time and access to do the work they are being asked to do.
When investigating the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises to investigate. Both later said they could not draw confident conclusions, due in part to scope and timing limitations.
A similar issue occurred during pre release testing for GPT-6 Astra, which OpenAI has touted as its most aligned model yet. According to Apollo Research’s contribution to the model card, the firm was given only three days to test Astra, which made it difficult to draw firm conclusions.
Apollo wrote in its evaluation that given the higher rates of evaluation awareness and the limited evaluation window, low rates of misbehavior did not provide substantial evidence about the model’s alignment or misalignment.
That track record leaves evaluators with a basic question: Why should this time be different?
Gleave put it plainly.
“It’s certainly possible that Dario and Sam just had a change of heart, and they’re going to be very open about this,” Gleave said. “But the intellectual property of these companies is so incredibly valuable to them, and I think they’re going to, by default, be very careful about what can be shared.”
The Case for Law, Not Goodwill
Several researchers who spoke to TechCrunch called for a transparent framework that all parties agree to publicly. Part of that framework, according to Steidley, should involve standards for what kinds of auditors companies can rely on. Without such standards, companies could sidestep scrutiny by shopping for evaluators that are either unqualified or uninterested in assessing the most concerning risks.
Henry Papadatos, executive director of Safer AI, says the deeper problem is that even a public framework built on voluntary measures always depends on a company’s goodwill.
“Ideally, we would have good regulation mandating this, because then companies cannot change their mind tomorrow if they have a big PR crisis,” Papadatos told TechCrunch. He noted that regulation is also a good means of pushing all companies to adhere to the rules, not only the most willing.
Papadatos delivered the sharpest line of the entire debate.
“You cannot have it both ways, having zero accountability externally, and then say, ‘I’ll just have my own flexible rules.'”
Not Everyone Is On Board
So far, Meta, SpaceXAI, and Google DeepMind have not committed to embedding third party evaluators. DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models. Google, OpenAI, and Anthropic have also privately been discussing AI safety plans for weeks.
Some laws are already forming around the idea of third party evaluators. California’s SB 53, signed into law last year, requires large frontier AI developers to publish safety frameworks and report critical safety incidents. A new law, SB 813, signed this month, creates a framework for state recognized independent verification organizations with expertise in assessing AI risks.
In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing, and to report serious incidents. The EU AI Office can also conduct its own evaluations and appoint independent experts.
For now, the law remains less expansive than what Amodei is proposing. That leaves frontier labs largely responsible for deciding how much independent scrutiny they will submit to.
The Bottom Line
Voluntary self regulation is better than nothing. But a watchdog that is hired, housed, and contractually constrained by the entity it is supposed to watch is not a watchdog in any meaningful sense. It is a vendor with a clipboard.
Amodei’s proposal is a genuine step forward in ambition. Altman’s quick endorsement suggests the industry recognizes that public trust is now a competitive asset. But ambition is not architecture, and an announcement is not a system.
The evaluators have told us exactly what they need: real access to training checkpoints, real time to investigate, real freedom to publish, and real legal backing so that the rules do not evaporate the moment a company faces a PR crisis.
Until those four conditions are met, embedded evaluation will remain what it has always been. A promise made in good times, subject to revision in bad ones.
The question is not whether AI companies will allow independent evaluators inside. The question is whether they will allow them to matter.
TechTrib.com is a leading technology news platform providing comprehensive coverage and analysis of tech news, cybersecurity, artificial intelligence, and emerging technology. Visit techtrib.com.
Contact Information: Email: news@techtrib.com or for adverts placement adverts@techtrib.com