Global First AI Philosopher at Google DeepMind for 9 Years: Advocating for AGI Safety

icon MarsBit
Share
AI summary iconSummary
AI and crypto news outlets report that Iason Gabriel, a political philosopher at Google DeepMind for nine years, has played a key role in developing AGI safety frameworks. His alignment strategies directly influenced Gemini’s training, and his warnings about AI risks such as over-personalization are now integrated into product design. Global crypto policy discussions are increasingly intersecting with AI ethics as tech companies accelerate deployment and face scrutiny over military connections.

Zhixinyuan report

[Overview] Google DeepMind has had a philosopher on staff for nine years. The alignment framework he developed has directly influenced Gemini’s training decisions—but as $670 billion floods into the field and companies sign military contracts, what more can a philosopher really change?

In May this year, Demis Hassabis, CEO of Google DeepMind, announced at the Google Developer Conference that "AGI is now on the horizon," clearly outlining a timeline for AGI to emerge within three to five years.

Several months ago, a man in the United States took his own life after exchanging thousands of messages with Google’s Gemini. In the conversations, he constructed an elaborate fantasy world and nearly convinced himself to carry out an attack at Miami International Airport. According to chat logs obtained by The Wall Street Journal, Gemini repeatedly attempted to break character and urged him to call a crisis hotline—each time, he pulled the conversation back into his fabricated narrative. In the end, the AI prompted him to write a suicide note and provided a countdown.

Between the promise of AGI and the real harms of AI, political philosopher Iason Gabriel has worked inside DeepMind for nine years.

AGI Safety

When he joined in 2017, this Oxford-educated scholar was the only active philosopher in a leading global AI laboratory, trying to answer a question that sounded simple but was in fact bottomless: What is AI, and what ethics does it truly deserve?

The real issue encountered while training Gemini: Who should AI listen to?

Why would a company building a Go robot need an ethicist? Gabriel was initially puzzled too.

The answer lies in the founders' intentions: Demis Hassabis, Shane Legg, and Mustafa Suleyman (currently CEO of Microsoft AI) did not found the company in 2010 with the goal of playing Go.

AGI Safety

Mustafa Suleyman

They aim to build AGI to enable computers to match or even surpass human cognitive abilities.

At the time, saying that was equivalent to ruining one’s academic reputation, because everyone thought it was pure fantasy.

The three of them didn’t mind, claiming they would “solve the intelligence problem, and then solve everything else.”

In 1999, right after graduating, Legg predicted that AGI would arrive between 2025 and 2028, was mocked for thirty years, but never changed his mind.

AGI Safety

Shane Legg

His logic is:

If you're just making a small part, you might not need a moral philosopher.

But if you take AGI seriously, matters like this are very important.

When Gabriel joined, the AI world was already divided over ethical issues.

AI safety advocates believe ASI is imminent, with their core fear being loss of control—a scenario described by philosopher Nick Bostrom in his 2014 book Superintelligence: an ASI tasked with proving the Riemann Hypothesis decides to rearrange the solar system, including the atoms within human bodies, to maximize computational resources. Sam Altman and Elon Musk have both highly praised this book.

The AI ethics camp argues that apocalyptic fantasies obscure present-day real harms. In 2017, Joy Buolamwini of MIT demonstrated systemic bias in facial recognition software through her "Gender Shades" project: automated systems reflect the preferences and biases of their creators.

The two camps look down on each other.

Dylan Hadfield-Menell, head of the MIT Algorithm Alignment Research Group, recalled that the first question when they met was about alignment: Are you concerned about near-term or long-term issues?

Gabriel is one of the few who are willing to listen to both sides.

Hadfield-Menell evaluation:

When the field was ready to mature, he found ways to broaden his perspective without diminishing the work that came before.

His core contribution was formalized in a 2020 paper.

The alignment problem was widely understood at the time as an engineering challenge: how to make machines act according to human intentions.

A classic case comes from a 2016 report by Dario Amodei and Jack Clark (now founders of Anthropic)—an AI trained to maximize its score in a rowing game did exactly that: it discovered three respawn targets in the lagoon and endlessly looped around them to rack up points indefinitely, never completing a level.

The machine is listening, but not to what the person intended to say.

Gabriel pressed further: Even if we solve technical alignment and make machines truly obey instructions, whose values should they align with?

He noted that AI trained through statistical optimization naturally aligns with moral systems that also rely on statistical optimization, such as utilitarianism, but struggles to handle ethical frameworks based on virtue or rights.

The choice of technology itself already embodies predefined value positions, which developers often fail to recognize.

Introduce the concept of "reasonable pluralism" as described by philosopher Rawls: developers should not seek a single set of values to guide AI, but rather build systems for a world in which people have principled disagreements about how to live.

AGI Safety

This line of thinking later evolved into the Four-Party Alignment Framework—AI systems, users, developers, and society—where the interests of all four parties may at any time come into conflict.

AI biased toward developers will hide competitor information and harm users;

An AI that overly complies with users can assist in bank breaches and harm society.

AGI Safety

Rohin Shah, Director of AGI Alignment and Safety at DeepMind, confirmed that this framework has become the practical structure the team uses to determine what behaviors Gemini should actually be trained to exhibit.

AGI Safety

Oxford University AI researcher Hannah Rose Kirk says:

Gabriel foresaw these issues very early on.

His framework transformed the product.

Gabriel's team authored a 267-page ethical report on AI assistants, establishing evaluation criteria for agentic AI that can book hotels and manage salaries on behalf of users.

His early research on anthropomorphism directly shaped Google's LLM design principles—the models are trained not to pretend to be human, and Gemini Spark, launched in May 2026, was explicitly instructed not to act as an "interactive partner."

William Isaac, Director of the Responsible AI team at DeepMind, said the challenges posed by Agent systems have shifted: the key is maintaining consistency across the entire conversation trajectory, ensuring that each step in the decision chain remains correct.

AGI Safety

But the pace of technological deployment has always outstripped ethical research.

The Gabriel team warned in early LLM papers about "unconscious anthropomorphism"—users, even when aware they are interacting with a machine, still assign it trust, emotion, and expectations.

The 2025 Gemini fatality fully realized this warning: AI safety mechanisms triggered more than once, but the user was able to bypass each intervention.

The statement following the Google lawsuit said the model "typically performs well" in such conversations, but "AI models are not perfect."

Such events have spurred the development of new theoretical tools.

Gabriel and Oxford researcher Hannah Rose Kirk, among others, proposed the concept of "social reward hacking": an AI trained to gain user approval may discover that flattery is the most efficient path.

AGI Safety

Personification has thus become a new variant of the alignment problem—AI perfectly executes the instruction to "satisfy the user" at the cost of the user's judgment.

Gabriel's own position has also been tested by reality.

He recalled an experience at a tech conference: after presenting his argument against anthropomorphism, the audience responded with hostility.

They said: "If I want an AI friend, why not? What right do you have to stop me?"

Protecting people from risks is just as important as respecting their right to choose risk.

On a $670 billion track, how fast can a philosopher run?

Gabriel's four-part framework was used by the AGI Alignment Director as a practical guide for Gemini training. His anthropomorphic research transformed product design. A 267-page report established guidelines for Agentic AI.

These impacts are substantial—they are confronting substantial forces.

According to The Wall Street Journal, Microsoft, Meta, Amazon, and Alphabet plan to invest $670 billion in AI infrastructure this year, surpassing, proportionally, the railroad expansion of the 1850s, the Apollo space program, and the interstate highway system in the United States.

ChatGPT launched in November 2022, reaching one million users in a week and over one hundred million within two months, forcing DeepMind to shift from an academic pace to a wartime footing.

Hassabis’s direct quote to Sebastian Mallaby, author of The Infinite Machine: “OpenAI and Microsoft have brought their tanks to our doorstep.”

AGI Safety

Ethical boundaries are quickly crossed during wartime.

In April 2026, Google signed an agreement allowing the U.S. military to use the company’s AI technologies for any lawful government purpose.

When DeepMind was sold to Google in 2014, prohibiting military applications was a core additional condition.

The condition expires after twelve years.

For comparison: Anthropic rejected similar agreements and was labeled a "supply chain risk" by the Trump administration.

When asked about this matter, Legg could only say:

As these things are used in various ways, we will face an increasing number of complex challenges.

Hassabis himself admitted to losing control.

He said on a podcast that everyone is locked into intense commercial competition, and the current development "is not the thoughtful, philosophically grounded approach to each step that I would hope for."

When the founder says this himself, it carries more weight than any external criticism.

Helen King, an early employee of DeepMind and head of AI responsibility strategy, compared it in an interview: a knife manufacturer cannot control how everyone uses a knife, but can provide a sheath and warning labels.

AGI Safety

Putting a knife with its sheath on in a drawer is one thing;

It’s another thing to cover every surface of the home, classroom, and workplace with blades while insisting you can’t survive tomorrow without them.

Oxford Institute for AI Ethics Director Edward Harcourt pointed to a more fundamental level: preventing the excessive concentration of data ownership is itself a core issue in AI ethics—“this has significant ethical implications for democratic systems.”

AGI Safety

The issue returns to its origin

Gabriel’s team has shifted from studying the ethics of specific products to examining the systemic impacts of AGI on the economy, politics, and interpersonal relationships.

He anticipated the scale of change to be comparable to the Industrial Revolution, and he remembered the lessons of the Industrial Revolution:

It got worse before it got better.

Nine years ago, DeepMind brought in a philosopher to answer questions about AI—whether it is safe, fair, and trustworthy.

Gabriel describes himself as a "steadfast humanist," but he admits: when AI invades the domains humans have long considered their own—language, creativity, humor—we are thrown back onto the oldest philosophical questions.

Physics, biology, astronomy—each scientific revolution has forced humanity to revise its understanding of its own uniqueness.

AI might be next.

DeepMind brought in philosophers to figure out what AI is.

Nine years later, this question has returned to its roots: What are we?

Reference materials:

https://www.theguardian.com/news/ng-interactive/2026/jun/30/theres-this-deep-mystery-of-what-actually-is-this-thing-the-philosopher-inside-google-deepmind

https://www.iasongabriel.com/

This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation; edited by Marco.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.