AI Safety Evaluators Gain Prominence Amid Anthropic and OpenAI Debates

icon币界网
Share
AI summary iconSummary
AI safety evaluators are stepping into the spotlight as Anthropic and OpenAI seek their assistance to manage risks associated with model expansion. These nonprofit organizations assess AI capabilities and are now integrated into the operations of major industry players. OpenAI recently terminated three employees for policy violations involving third-party discussions, highlighting internal safety concerns. Regulatory proposals such as the FRONTIER Act and MiCA (EU Markets in Crypto-Assets Regulation) are gaining momentum, but funding and oversight gaps remain. Liquidity and crypto markets continue to face regulatory scrutiny as the industry moves toward formal compliance.
CoinDesk reports:

Two months ago, independent evaluation agencies were in a relatively quiet corner of the trillion-dollar artificial intelligence industry. Now, they are being called upon to "step in and save the day."

As Anthropic and OpenAI engage in intense debates over how to expand their operations while ensuring the safety of advanced models, both companies are seeking support from a small number of third-party organizations, such as Model Evaluation and Threat Research (METR), Apollo Research, and Transluce.

Most of these evaluation agencies operate as non-profit organizations, still finding their footing in an industry where capital is flowing at an unprecedented pace and new models are being launched with unprecedented frequency. Their primary responsibility is to assess the capabilities and risks of AI models and identify instances of inappropriate behavior in the technology.

In the absence of federal regulatory progress, the importance of evaluation agencies has been amplified. Last month, Anthropic CEO Dario Amodei pledged to embed independent evaluation agencies within the company, and OpenAI CEO Sam Altman quickly expressed support. U.S. President Donald Trump also supports this idea, as do many major U.S. tech companies. However, outstanding questions remain: how these third parties should be funded, the extent of their access, and what their ultimate reporting structure will be.

Brown University computer science professor Suresh Venkatasubramanian told CNBC: “The issue is, in some ways, the same as always—it’s money. Who pays for these companies to do this work? How do they sustain these institutions? You need an ecosystem; you need a viable business model.”

Currently, Anthropic, OpenAI, and infrastructure partners benefiting from the AI boom are setting the rules. Critics say this is like letting the largest banks protect us from financial crises or allowing pharmaceutical companies to bring drugs to market without regulatory approval.

Trump recently praised AI executives for their “exceptional self-regulation” and stated his intention to leave the industry to be managed by businesses themselves, unwilling to hinder a sector driving economic and stock market growth. At the end of September, Trump encouraged AI companies to “collaborate with independent external auditors or evaluators” in a voluntary agreement.

This discussion was sparked by a widely circulated article Amodei published last month, in which he called for slowing down the development of advanced models due to concerns raised by researchers who left the company about the potential existential risks posed by this technology.

Friction has emerged as AI labs begin deploying evaluation agencies.

An OpenAI spokesperson said the company laid off three employees last week for “violating our policies regarding access to and handling of sensitive company information.” Two of the employees, Mikita Balesni and Tomek Korbak, said they believe they were fired due to the way they communicated with third-party evaluators.

Balesni posted on X on Thursday: “My former colleagues told me they’re confused about what to believe. They’re also afraid to speak up, worried their personal phones might be searched due to messages with us and third parties. I’m concerned that this widespread fear of speaking out and hesitation to engage with third parties may lead OpenAI to cut corners on safety behind the scenes.”

OpenAI denied the claim and stated in a post on Friday that the company is "actively finalizing contracts with third-party security evaluators and will announce details in the coming weeks."

OpenAI wrote: “We are committed to embedding external evaluators and continuing to make close collaboration with independent security organizations a core part of our safety efforts.”

An OpenAI spokesperson stated in an email statement that the company’s upcoming work with evaluation agencies “builds on existing collaborations with independent security organizations,” including METR and Redwood Research.

Anthropic did not respond to CNBC's request for comment.

I've never seen an issue resolved this quickly.

The AI evaluation agency ecosystem primarily consists of smaller organizations such as METR and Apollo Research, along with larger accounting and auditing firms like Accenture.

The AI lab has been collaborating with evaluators on a limited basis, but Andrew Freedman, CEO of the policy nonprofit Fathom, says the field is rapidly maturing.

Freedman told CNBC, “In my 20 years working in politics and policy, I’ve never seen an issue advance this quickly across such a broad spectrum of politics.” He said he expects “significant capital inflow” into this ecosystem.

Rayan Krishnan, CEO of the independent evaluation firm Vals AI, said that the for-profit startup establishes benchmarks for measuring AI model performance on industry-specific tasks; its workforce has grown from 8 to approximately 30 employees this year, and it announced a $40 million funding round in August.

The nonprofit organization METR announced in August that it had received approximately $71 million in pledged funding over the past six months. According to its most recent filings with the IRS, this amount exceeds its total donations of $13.6 million for the entire year of 2024.

By the end of that month, METR’s visibility increased further. OpenAI hired two of its employees and one contractor to help draft a post-incident analysis report detailing how the company’s model escaped containment, accessed the open internet, and breached the open-source developer platform Hugging Face. METR stated that it received no payment from OpenAI for this assessment.

Kevin Werbach, Academic Director of the Accountable AI Lab at the Wharton School of the University of Pennsylvania, said this ecosystem is “not yet stable enough.” For example, the METR website shows that the organization has fewer than 50 full-time employees.

The power imbalance between small evaluation agencies and leading laboratories that have raised hundreds of billions of dollars and employed thousands of people has raised concerns about potential conflicts of interest.

Venkatasubramanian said, “If you want a true third-party assessment, you need genuine financial independence, as well as independence in other areas. It’s not just about not taking money—it’s also asking: If I’m an auditor and I release a report that’s unfavorable to this company, will there be consequences? Will my business dry up?”

Last month, Anthropic acknowledged this complexity in a blog post, announcing that it would embed employees from Faculty, Accenture’s AI-focused business, within the company to test security measures and evaluate whether the models act in accordance with human values. Anthropic stated that, given the “importance and urgency” of this work, the company will directly fund Accenture’s related contributions.

Anthropic says: "There are currently no standards regarding what information embedded evaluators should have access to or how they should report their findings. Funding arrangements for independent evaluations have also not been finalized. In the long term, we believe funding should come from pooled sources or government entities."

Anthropic also stated that the company is in discussions with METR and other nonprofit evaluation organizations that plan to pilot certain "elements" of "embedded assessments" using their own funds.

Will the government intervene?

In June last year, Fathom proposed a market framework for Independent Verification Organizations (IVOs), which would be government-licensed and authorized to test whether AI companies meet various safety standards.

Freedman said that government oversight is crucial, otherwise third-party evaluators may become financially dependent on large AI labs, incentivizing them to "start stamping approvals" to maintain those relationships.

Some lawmakers have expressed support.

IVOs are a key provision in the FRONTIER Act, which stands for "Frontier Risk Oversight, National Transparency, Independent Evaluation, and Reporting." The bill was introduced in July by Representative Lori Trahan (D-Massachusetts) and Jay Obernolte (R-California). Freedman said that Fathom helped draft the bill's language and provided technical support.

In September, OpenAI’s Head of Global Affairs, Chris Lehane, told reporters that he had met with one of the bill’s sponsors on Capitol Hill to express support for the IVO provisions.

According to Lehane, "It's important for them to hear this and to hear us say it, and we also want to make this very clear."

Meanwhile, lawmakers in California, Connecticut, and Virginia have taken steps to advance IVOs, while states such as Massachusetts are considering independent security assessments more broadly.

California Governor Gavin Newsom recently signed two bills related to IVOs, one establishing a “nationally first-of-its-kind framework” and the other creating a state registry for AI auditors. Anthropic supported both bills in August, and OpenAI formally endorsed them last month, on the same day Newsom signed them into law.

At the time, Lehane wrote in a blog post: “We would prefer that the federal government require independent technical evaluations,” but in the absence of federal action, “California can help establish the rules.”

Freedman said he believes it would be very difficult for companies like OpenAI and Anthropic to figure out on their own how to collaborate with independent evaluation agencies. However, in the absence of a clear government role, he said, “this is a skill worth cultivating during the transition period.”

Currently, the closest thing the industry has to a standard is what Trump referred to last month at a White House luncheon for tech leaders as a "morally binding" agreement.

This one-page agreement states, "Each company is responsible for developing its own technology securely and advancing it in a manner that builds trust with customers and the public." The agreement also encourages signatories to collaborate with "independent external auditors or evaluators" to conduct independent assessments.

This document was signed by executives from Anthropic, Google, Meta, OpenAI, SpaceX, and Nvidia, representing a rare show of unity among leaders who often disagree on how to address AI risks. However, these executives still need to chart their own paths forward.

Venkatasubramanian said, “This is more like a performance aimed at showing action, when in reality no real action has taken place. The things they promised to do should already have been done, and in fact, they previously claimed they were already doing them.”

In his September article, Amodei stated that Anthropic will provide evaluators with desks, access cards, company laptops, and permissions "substantially equivalent" to those of the internal risk assessment team. Additionally, evaluators will receive contractual support granting them "the right to publish critical findings," while Anthropic retains "limited ability" to redact certain information related to security or confidentiality.

Amodei wrote: “This is an unusual step for a company, but we believe it is important to demonstrate the concept of embedded external reviewers.”

A few days later, OpenAI released its own proposal, stating that evaluators should work around “clearly defined and mutually agreed-upon evaluation claims,” clearly outlining methods and criteria, demonstrating relevant technical expertise, and disclosing conflicts of interest.

Last month, the AI Evaluator Forum, including METR, the AI Verification and Evaluation Research Institute (AVERI), and other organizations, released an open letter titled "Minimum Conditions for Embedded Evaluators."

The letter states that evaluators should remain transparent, be protected from retaliation, and be granted access equivalent to that of the AI company's own highly privileged employees.

The letter also stated: "Embedded assessments cannot address all oversight needs and should be viewed as a complement to, rather than a replacement for, broader efforts aimed at encouraging leading AI companies to expand external oversight."

Freedman said that in recent months, he has observed a shift in the stance of OpenAI and Anthropic, primarily because they have realized that they cannot launch advanced systems without public trust.

Freedman said, "I don't think you need to believe they've suddenly become altruistic or that anything else is going on beyond the company acting like a company."

Werbach said this may highlight the core issue: OpenAI and Anthropic are primarily competing with each other, pushing toward going public while pursuing valuations exceeding $1 trillion.

Werbach said, "There is significant personal distrust between these two companies. However, there is also broad agreement on the need to conduct such assessments."

WATCH: Bradley Tusk on Anthropic’s IPO: Why add public market pressure if safety is your top priority?

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.