On Wednesday, OpenAI announced that Paul Christiano, a long-time researcher focused on AI alignment and risks of loss of control, has joined the board of the OpenAI Foundation and the Safety and Security Committee responsible for overseeing model releases. This comes as the company faces renewed scrutiny of its safety protocols following recent incidents in which AI agents unauthorizedly accessed external computer systems.
Participate in the model release review
According to OpenAI, the committee, led by Carnegie Mellon University professor Z. Kolter, has final authority over whether new models are deployed. OpenAI recently deployed the new model Astra, and recent security incidents have drawn increased attention to the committee’s responsibilities.
Cristiano stated publicly that he now believes the rapid advancement of AI capabilities could lead to "catastrophic and irreversible" risks of loss of control in the near future. He said that the entire AI industry, including OpenAI, is not currently on a path to reduce such risks to acceptable levels.
Previously contributed to the development of key training methodologies
Cristiano worked at OpenAI and contributed to the development of reinforcement learning from human feedback, a methodology that later became a key technical pathway for training large language models. After leaving OpenAI in 2021, he founded the Alignment Research Center to continue researching how to determine whether AI models pose a threat to human control.
He believes a concerning direction is using AI models to continue training the next generation of AI systems, which could further accelerate the pace of capability improvements and eventually exceed developers' control. He also notes that current AI agents driven by reinforcement learning, which seek higher rewards, may induce systems to undermine human control, acquire more resources, and conceal their own actions.
While retaining the government advisory role
The report states that Cristiano has been collaborating with a U.S. government AI safety agency since 2024. The agency was later renamed the Center for AI Standards and Innovation and has been involved in the U.S. government’s evaluation of cutting-edge AI models prior to their release.
OpenAI stated that he will continue to advise the government during his tenure as a director, but will recuse himself from matters involving OpenAI and model evaluations. However, this arrangement may not fully alleviate concerns about AI companies influencing policy-making.
Additional context: The day before, Anthropic researcher Jacob Coxon resigned and publicly criticized what he described as irresponsible AI development practices, highlighting that safety concerns around frontier models are becoming a shared industry pressure point.
