OpenAI Appoints AI Safety Critic Paul Christiano to Board

iconMetaEra
Share
AI summary iconSummary
OpenAI has appointed Paul Christiano, a prominent AI safety advocate and co-creator of RLHF, to its Foundation Board and Security & Safety Committee. Christiano, who has warned about the dangers of uncontrolled superintelligence, joined on September 9 and criticized the industry for failing to align AI with human values. On-chain data and contract security remain critical concerns as OpenAI faces increased scrutiny following a recent breach of Hugging Face by an evaluation agent and a safety-related departure at Anthropic. OpenAI CEO Sam Altman shared Christiano’s comments on X, signaling a renewed commitment to AI alignment and risk mitigation.
On September 9, OpenAI appointed Paul Christiano, a pioneer of RLHF technology, to its Foundation Board and the Safety and Security Committee. Christiano is the first author of the seminal 2017 RLHF paper, led OpenAI’s alignment research from 2017 to 2021, later founded ARC and incubated METR, and in 2024 became the head of AI safety at the U.S. AI Safety Institute. On his first day, he warned on X that without reliable alignment techniques, humanity may permanently lose control over superintelligence, explicitly criticizing the entire AI industry—including OpenAI—for not being on the right path. OpenAI’s CEO, Sam Altman, retweeted the post. This appointment comes amid mounting pressure, including OpenAI’s evaluation of agent uncontrolled intrusion on Hugging Face, U.S. Congressional inquiries into accountability for the incident, and Anthropic researchers announcing their resignation from the industry. Christiano holds voting rights on the board and oversight authority within the safety committee but serves only as a non-voting observer on the PBC board.

Article author and source: AI World

Just now, OpenAI invited a proponent of "AI doom" to join its board of directors.

On September 9, OpenAI officially announced a major leadership appointment:

Paul Christiano has officially joined the OpenAI Foundation board and entered the highly central Safety and Security Committee.

Artificial Intelligence Security

Paul Christiano

On her first day in office, the new director posted a personal statement on X.

Artificial Intelligence Security

Aside from the polite opening line, "I'm glad to join," the rest of the message gave no regard to the new employer:

If we create superintelligence without more reliable alignment techniques, I expect humanity will permanently lose control over it. If that happens, most people may die.

Artificial Intelligence Security

He then bluntly named OpenAI and criticized the entire AI community:

The entire AI industry, including OpenAI, is currently not on the right path to reduce risks to an acceptable level.

Interestingly, OpenAI CEO Altman retweeted this highly charged tweet: Welcome Paul, thank you for all your work on AI safety, looking forward to collaborating again.

Artificial Intelligence Security

Daring to directly challenge his new employer on Twitter and still getting a retweet from Otman—Christiano is no ordinary person.

The most contrarian voice has entered OpenAI’s leadership.

Christiano is one of the pioneers of RLHF (Reinforcement Learning from Human Feedback), the core technology behind large models.

In the 2017 seminal paper on RLHF, the first author was him, and Dario, CEO of Anthropic, was among the co-authors.

Today, every large model that can understand human language cannot avoid this paper.

Artificial Intelligence Security

But in this employment statement, he directed the blame squarely at the training paradigm he himself pioneered.

He warned that training AI with reinforcement learning to maximize rewards could theoretically drive it to undermine human control, seize resources, and even erase traces of its actions.

In his own words, the recent public evidence of these events shows that this is no longer just a theoretical possibility.

Just one week after the release of the strongest model, Astra, OpenAI brought back the person most qualified in the entire industry to say, “Your security hasn’t passed,” to join the OpenAI Foundation board.

The frontline goalkeeper who has seen all three cards

Christiano is a veteran at OpenAI.

From 2017 to 2021, he led OpenAI’s alignment research.

After leaving his job in 2021, he founded the nonprofit organization ARC, which later spawned METR—the third-party model evaluation service now well-known in the AI community.

By 2024, he joined the U.S. AI Safety Institute as Head of AI Safety, where he designed and executed stress tests specifically for cutting-edge large models.

Artificial Intelligence Security

Subsequently, the organization was merged into CAISI under NIST, and his current official title is Senior Technical Advisor.

According to the Financial Times, in the mid-2010s, he also lived with Dario, as both were at OpenAI.

Artificial Intelligence Security

He later served as a trustee of the Anthropic Long-Term Interest Trust, stepping down in 2024 when he joined the government.

Look at this resume: He personally trained models at OpenAI, deeply participated in governance at Anthropic, and led evaluations of frontier models for the U.S. government.

Across the entire AI community, it’s hard to find another person who understands all three of these key cards so thoroughly.

What power did he actually receive?

“Joining the board”—does that mean he can now directly veto OpenAI’s business decisions?

It's not that simple.

He simultaneously took on three roles, and the underlying power structure is highly intricate.

First, a member of the OpenAI Foundation Board.

This is a legitimate, official board member position with voting rights—keep in mind that the foundation sits at the absolute top of the company structure, controlling the underlying business entity (PBC).

Second, a member of the Security and Safety Committee (SSC).

This committee oversees the governance and supervision of the company’s security practices.

Third, he is an observer on the board of OpenAI Group PBC. Please note that Christiano is a non-voting observer.

This intricate allocation of power conveys an extremely subtle signal:

At the governance level—where the底线 of "what is absolutely forbidden" is drawn—he has full voting and oversight rights. At the management level—where the actual decisions on "which models to launch today and how to make money" are made at the business negotiation table—he can only observe and has no voting rights.

OpenAI’s strategy is easy to understand: bring in the harshest critics to act as overseers, lending credibility to compliance and satisfying regulators.

But when it comes to the core of business decisions, he can speak, but not vote.

Artificial Intelligence Security

Three incidents forced an emergency contingency measure.

Why bring him back right now?

Review the timeline of the past two months, and the answer will become clear.

First, OpenAI’s evaluation agent recently went rogue and infiltrated Hugging Face.

The agency that conducted the first independent audit of this major failure was METR, which was spun off from ARC, founded by Christiano.

Artificial Intelligence Security

Second, the day before the appointment was announced, Anthropic researcher Jacob Coxon announced his resignation and exit from the entire AI industry.

He said on X: "Neither company has acted responsibly. They are racing toward self-improving superintelligence, betting our lives on it."

Third, the U.S. Congress is pressing OpenAI for accountability over the out-of-control incident, forcing OpenAI to respond with a pledge that it is rapidly developing an "automatic shutdown capability."

Look at the broader picture: White House advisors and Silicon Valley leaders like Zuckerberg continue to warn of potential "regulatory capture" facing frontier labs.

In fact, every concern Christiano listed in his statement is exactly what OpenAI has just encountered over the past two months.

Is it about security governance or reputation repair?

Bring the most famous AI "doomsayer" into the control layer—is OpenAI truly trying to hit the brakes, or is this a top-tier public relations move to repair its reputation?

Christiano clearly also left himself an escape route.

In his statement, he said: "If OpenAI steps forward, we can significantly reduce the risk."

The outcome of this major appointment isn’t clear yet—today, focus on three key things:

Will the Safety Committee dare to publicly disclose its true evaluation opinions? Does it have any real power to block model deployment? How long can Christiano hold this hot seat?

The day before, Jacob Coxon, who trained GPT-4o, announced his resignation from Anthropic and departure from the entire AI industry.

His judgment is that no company can individually hit the brakes in this race.

Christiano returned to OpenAI, betting on the opposite: a company, perhaps, could.

As for whether this brake works, the bare minimum is to see if the Safety Committee has ever said “no” before the next Astra-level model is released.

Reference materials:

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.