AI Safety2026-09-17TechCrunch AI

Anthropic, OpenAI Want Embedded Safety Evaluators

Anthropic and OpenAI are exploring a new approach to AI safety: placing independent evaluators inside their own labs. The idea is to give outside experts early access to models, training pipelines, and deployment decisions so risks can be spotted before products reach the public. Researchers say this level of access is unprecedented. It could let evaluators influence design choices, test dangerous capabilities, and flag problems while fixes are still cheap. Supporters argue internal scrutiny can catch issues faster than external audits that happen after a model ships. But critics warn that embedded evaluators may face pressure to soften findings. If their funding, career prospects, or access depend on the labs, can they truly challenge commercial priorities? Independence, transparency, and clear reporting lines would be essential. Evaluators also need the ability to publish concerns, escalate to regulators, and walk away without retaliation. The debate reflects a broader question in AI governance: can voluntary internal oversight work at companies racing to build powerful systems? Some experts see embedded review as a useful stopgap, not a replacement for regulation. They want mandatory external audits, safety standards, and legal accountability. Others believe close collaboration between labs and evaluators can improve safety culture and speed up safeguards. For now, Anthropic and OpenAI have signaled interest, but details remain unclear. The key test will be whether embedded evaluators get real power or become a public-relations layer. If they can challenge leaders, shape releases, and disclose risks, the model could meaningfully improve accountability. If not, it may create an illusion of oversight while major decisions stay with the labs.

Related news