OpenAI announced on Wednesday that Paul Cristiano, who has long studied the alignment and out-of-control risks of AI, has joined the OpenAI Foundation board of directors and will also be part of the Security and Safety Committee responsible for overseeing model releases. At this time, the company is once again facing a review of its security procedures due to several recent incidents where AI agents exceeded their authority by accessing external computer systems.
Will participate in model release review.
According to OpenAI, this committee is led by Professor Z from Carnegie Mellon University, who is in charge of token issuance Kolter. They have the final decision-making power regarding whether a new model will be launched. OpenAI just deployed a new model Astra last week, and recent security incidents have also drawn more attention to the responsibilities of this committee.
Cristiano stated in a public statement that he now believes that AI has the potential to improve rapidly, which may lead to "catastrophic and irreversible" risks of loss of control in the near future. He mentioned that the entire AI industry, including OpenAI, is not on track to reduce such risks to an acceptable level.
Involved in the development of key training methods
Cristiano once worked at OpenAI and was involved in promoting the development of reinforcement learning based on human feedback. This method later became an important technical approach for training large language models. After leaving OpenAI in 2021, he founded Alignment Research Center to continue researching how to determine whether AI models pose a threat to human control.
He believes that a direction worthy of vigilance is to continue training the next generation of AI systems using the AI model. This could further accelerate the rate of capability improvement, ultimately exceeding the control of developers. He also pointed out that the current use of reinforcement learning to drive AI agents to pursue higher rewards may induce the systems to weaken human control, strive for more resources, and conceal their own actions.
While also retaining the role of government advisor
The report mentioned that Cristiano has been cooperating with a U.S. government AI security agency since 2024. This agency was later renamed the Center for Artificial Intelligence Standards and Innovation and is involved in the U.S. government's work related to the pre-release evaluation of cutting-edge AI models.
OpenAI indicates that he will continue to provide advice to the government during his tenure as a director, but he will recuse himself from matters involving OpenAI and model evaluations. However, this arrangement may not necessarily dispel concerns from the outside world regarding the impact of AI company on policy-making.
Additional information:Just the day before, Researcher Anthropic Jacob Coxon resigned and publicly criticized what he referred to as the irresponsible development practices of AI, indicating that security issues with cutting-edge models are becoming a common concern for the industry.










