Edited by Panda
A few days ago, Google lost two key personnel in succession.
On June 18, Noam Shazeer, one of the co-authors of the Transformer paper, announced on X that he was leaving to join OpenAI. Two days later, John Jumper, the 2024 Nobel Prize winner in Chemistry and head of the AlphaFold team, also announced his departure from Google DeepMind to join Anthropic.
Two consecutive announcements sent shockwaves through the capital markets: Alphabet, Google’s parent company, saw its stock plunge more than 7% at one point, erasing over $300 billion in market value. Multiple analysis firms attributed this sell-off to “talent defections.” D.A. Davidson analyst Gil Luria bluntly stated that Shazeer’s move to OpenAI and Jumper’s departure to Anthropic—occurring in quick succession—have led the market to fear that Google is losing ground in the battle for AI talent.

Shazeer’s departure is particularly noteworthy—it’s his second time leaving Google.
In 2021, he left the company after becoming dissatisfied with its refusal to publicly release the chatbot he led the development of, founding Character.AI; in August 2024, Google paid approximately $2.7 billion for the technology license of Character.AI and brought him back to DeepMind, appointing him as Engineering Vice President for the Gemini project, co-leading it with Jeff Dean. Less than two years later, he left again—this time joining OpenAI, Google’s longtime rival.
Thus, all eight co-authors of the paper "Attention Is All You Need," published nine years ago, have left Google.
User Tyler Maran created a chart showing where each of them stands today, and the chart has been wildly shared on social media.

However, this chart may soon become outdated. Over the past two days, rumors have circulated that NVIDIA is quietly bringing on board the core team of Essential AI, including Ashish Vaswani, co-founder and CEO of Essential AI and one of the authors of the Transformer paper. As of publication, neither NVIDIA nor Essential AI has issued an official response to these reports.
Seizing this opportunity, let’s take a comprehensive look at the past nine years of these eight individuals dubbed the "Fathers of Transformers," and where they truly are today.
It should be noted that the author order in the paper "Attention Is All You Need" is randomly arranged. The footnote clearly states that all authors contributed equally and the order was randomized; therefore, there is no such thing as a "first author" or "corresponding author." This article will introduce the eight authors in the original order listed in the paper.
"The Origin of Everything": Eight Google Employees Who Strayed from Their Jobs
To understand where they are today, we need to go back to 2017. At that time, the dominant approach in machine translation was recurrent neural networks (RNNs), which processed sentences word by word in sequence—like cars lining up one by one on a one-way street—making parallel computation impossible and rendering training slow and expensive.
Eight researchers from Google Brain decided to test a nearly audacious idea: discard recurrent structures entirely and keep only the "attention mechanism," allowing the model to view an entire sentence at once and determine for itself which words deserve the most focus. The paper’s title, “Attention Is All You Need,” is a play on The Beatles’ song “All You Need Is Love,” and has since become a popular template for many subsequent paper titles.

Author contributions statement, briefly detailing each person's specific contributions:

Jakob Uszkoreit was the first to propose replacing recurrent structures with self-attention and led the early validation of this idea;
Ashish Vaswani, along with Illia Polosukhin, designed and implemented the original Transformer model, contributing to nearly every aspect of the project;
Noam Shazeer proposed scaled dot-product attention, the multi-head attention mechanism, and parameter-free positional representations, and is another individual who personally handled nearly everything;
Niki Parmar designed, implemented, and debugged countless model variants in the original codebase and the subsequent tensor2tensor framework;
Llion Jones also experimented with numerous new model variants and was responsible for the initial codebase, inference efficiency optimizations, and visualization work;
Łukasz Kaiser and Aidan N. Gomez spent countless nights and days building the various modules of tensor2tensor, replacing the early codebase and significantly improving experimental results and research efficiency.
This description also indirectly reveals a detail: although the author order of the paper was randomized, Uszkoreit, Vaswani, Polosukhin, and Shazeer clearly took on more central roles in the architecture, while Parmar, Jones, Kaiser, and Gomez led the engineering implementation and system construction—this early distinction foreshadowed the differing paths and divergent strengths each of the eight would later pursue.
The name "Transformer" also has an interesting backstory. Uszkoreit liked the sound of the word, so the team internally began calling themselves "Team Transformer," and early design documents even featured illustrations of the six characters from the Transformers animated series on their covers.
Since its publication, the paper has been cited over 260,000 times and is one of the most cited papers of the 21st century.

Ashish Vaswani

Vaswani, born in 1986, is Indian. He earned a bachelor’s degree in computer science from Birla Institute of Technology, Mesra, India, in 2002, then moved to the United States to pursue his Ph.D. at the University of Southern California under David Chiang, focusing on statistical machine translation and neural network language modeling. After completing his doctorate, he worked as a computer scientist at the University of Southern California’s Information Sciences Institute for two years, and in 2016 officially joined Google Brain as a research scientist, where he worked until 2021.
According to the author contributions in the paper, Vaswani, together with Illia Polosukhin, designed and implemented the original Transformer model and was one of the key figures who "participated in nearly every aspect of the project."
After leaving Google, Vaswani co-founded Adept AI in 2021 with Niki Parmar, former OpenAI Vice President of Engineering David Luan, and others, serving as Chief Scientist with the goal of building "action models" capable of autonomously performing tasks within any software.

Adept once raised over $400 million, with a valuation of approximately $1 billion, but its product failed to materialize, and internal disagreements emerged within the team. Vaswani and Parmar departed early—Vaswani’s tenure as Chief Scientist at Adept ended in November 2022.
At the beginning of 2023, Vaswani and Parmar reunited to co-found Essential AI, with Vaswani serving as CEO. The company secured strategic investments from Google, NVIDIA, and AMD: its $8.3 million seed round was led by Thrive Capital, and its $56.5 million Series A round at the end of 2023 was led by March Capital, with participation from Google, NVIDIA, AMD, KB Investment, Franklin Templeton, and others. In early 2026, the company completed a $175 million Series B round led by Lightspeed Venture Partners, with Thrive Capital participating, achieving a $1 billion valuation and officially becoming a unicorn.
By the end of 2025, the company released its first open-source model series, Rnj-1 (named after the Indian mathematician Ramanujan).
However, over the past two days, the tide has turned. Reports indicate that NVIDIA is recruiting the core team of Essential AI, including Vaswani himself, who will contribute to the development of NVIDIA’s open-source model, Nemotron.
Sources reveal a rather practical reason: Essential AI’s fundraising is hitting a bottleneck, and luring Vaswani and his team away from AMD—NVIDIA’s competitor and an early strategic investor in Essential AI, which has long relied on AMD’s GPUs—is itself a worthwhile move. Several Essential AI researchers, including Alok Tripathy and Saurabh Srivastava, have already updated their LinkedIn profiles to show they have joined NVIDIA. However, as of now, neither NVIDIA nor Essential AI has officially confirmed the news.
Noam Shazeer

Shazeer was born in 1976 in Philadelphia and is an Orthodox Jew; his father, Dov Shazeer, is an engineer with a background in mathematics, and his sister was ordained as a rabbi by the Hebrew College. He demonstrated exceptional talent from a young age, earning a perfect gold medal as a member of the U.S. team at the International Mathematical Olympiad in 1994. He then attended Duke University to study mathematics and computer science, where he was awarded the Angier B. Duke Memorial Scholarship and received recognition in the Putnam Mathematical Competition.
In 2000, Shazeer joined Google, where his early breakthrough was fixing the spell-check feature in Google Search.
According to the author contributions in the Transformer paper, he proposed scaled dot-product attention, multi-head attention, and parameter-free positional representations, and was the person who "participated in nearly every detail" besides Vaswani and Polosukhin.
After co-authoring the Transformer paper in 2017, he and his colleague Daniel De Freitas developed the chatbot Meena, but Google chose not to publicly release it out of caution. In 2021, the two left to found Character.AI, which raised over $150 million from institutions including a16z and became a popular role-playing chat application.

In August 2024, the story took a turn: Google reached a licensing agreement with Character.AI, reportedly worth up to $2.7 billion. Shazeer and De Freitas returned to Google DeepMind with a small group of colleagues, and Shazeer was appointed Vice President of Engineering, co-leading the Gemini project alongside Jeff Dean and Oriol Vinyals. Since he personally held between 30% and 40% of Character.AI’s shares, the deal reportedly enabled him to cash out between $750 million and $1 billion. In 2026, he was elected to the National Academy of Engineering, with his career trajectory appearing to be at its peak.

But just a few months later, he chose to leave again, this time joining OpenAI, where he is reportedly set to lead a direction called "Architecture Research," just as OpenAI is ramping up hiring ahead of its IPO (the company secretly filed its S-1 with the U.S. Securities and Exchange Commission on June 8, with rumors valuing it as high as $852 billion).
OpenAI CEO Sam Altman made a rare public statement: “Since the very first day OpenAI was founded, he has been one of the people I most wanted to work with,” and added that this hiring had been “in the works for a full decade.”

For Google, this was an expensive failed repurchase: the person they brought back for $2.7 billion two years ago has now joined their top competitor, and this has become one of the direct triggers for Google’s stock plunge this week.
Niki Parmar

Parmar was born in Pune, India, and earned her undergraduate degree in Information Technology from the Pune Institute of Computer Technology. During her studies, she became interested in artificial intelligence and machine learning through online courses offered by Andrew Ng and Peter Norvig. She later moved to the United States to pursue a Master’s in Computer Science at the University of Southern California, where she worked with Professor Morteza Dehghani on applying machine learning methods to social science problems.
In 2015, Parmar joined Google Research as a software engineer, and in 2017 moved to Google Brain as a research software engineer—reportedly the youngest and the only researcher on the Google Brain team without a Ph.D.
According to the author contributions in the paper, she designed, implemented, and debugged countless model variants in both the initial codebase and the subsequent tensor2tensor framework. After the paper’s publication, she continued to extend the Transformer beyond language, contributing to research that applied self-attention mechanisms to image generation and computer vision.
In 2021, Parmar left Google and co-founded Adept AI with Ashish Vaswani, David Luan, and others, serving as Chief Technology Officer. She departed Adept early, much like Vaswani, and in early 2023, she co-founded Essential AI with Vaswani, continuing as a co-founder.

But she didn’t wait for Essential AI’s subsequent Series B funding or its unicorn status. At the end of 2024, Parmar quietly left Essential AI and joined Anthropic, publicly announcing the move in February 2025. She wrote on X: “As usual, today’s a good day to share: I joined Anthropic last December.”

She later contributed to the development of Claude 3.7 Sonnet—one of Anthropic’s most significant model releases to date. Today, she is a Member of Technical Staff at Anthropic, focusing on cutting-edge capability research and reinforcement learning.
Two former co-authors who were inseparable partners and co-founded two ventures together ultimately took vastly different paths: Parmar quietly stepped away over a year ago and quietly joined a leading laboratory, while Vaswani chose to continue pushing Essential AI forward—until this week, when he was finally picked up by a competitor.
Jakob Uszkoreit

Uszkoreit was born into a family of linguists. His father, Hans Uszkoreit, is a renowned computational linguist. Even his father was skeptical when the son proposed the hypothesis that "attention is all you need." Uszkoreit earned his Ph.D. from the Technical University of Berlin and later achieved the title of Distinguished Scientist at Google Brain.
According to the author contributions in the paper, it was Uszkoreit who first proposed replacing recurrent neural networks with self-attention mechanisms and led the early validation of this idea—the seed of this hypothesis was already planted in his 2016 paper co-authored with Ankur Parikh, Oscar Täckström, and Dipanjan Das, titled “Decomposable Attention Models.”

The name "Transformer" was chosen because he liked the sound of the word; the team internally refers to itself as "Team Transformer," and the early design document cover featured six characters from the Transformers animated series.
By the end of 2020, DeepMind’s AlphaFold2 demonstrated that Transformer-style models could solve the long-standing “holy grail” problem of protein folding. He increasingly realized that the reason deep learning had not yet truly transformed biology was not a lack of algorithms, but a lack of data. “It became almost a moral obligation,” he later recalled.
In 2021, he co-founded Inceptive with Rhiju Das, a Stanford University biochemistry professor and creator of the renowned RNA design game Eterna. Headquartered in Berkeley, the company maintains its research team in Berlin—where he lives—while also having employees across Zurich, London, Vancouver, and multiple cities on the U.S. East Coast. The company’s core idea is to reverse the traditional approach: instead of training models on existing data, it uses robots and human input to generate large-scale, novel RNA experimental data, which is then used to train the models.

Inceptive has raised approximately $1.2 billion from institutions including NVIDIA, a16z, Obvious Ventures, and Section 32. The latest development occurred this month: in early June, Alnylam Pharmaceuticals, a pioneer in RNA interference therapy, entered into a strategic partnership with Inceptive to leverage its foundational model to accelerate the design of siRNA drug candidates, with an upfront payment of $300 million and a potential total value of the collaboration reportedly reaching up to $2 billion. Uszkoreit stated in a declaration: “Most drug design still relies on trial and error—testing thousands of molecules in the hope that one will succeed. Inceptive takes a different approach: life follows extraordinarily complex rules that only AI can learn.”
Among the eight authors, he is the only one who fully switched to biotechnology, which precisely validates a prophecy from that paper: the potential of attention mechanisms extends far beyond machine translation.
Llion Jones

Jones is Welsh and graduated from the University of Birmingham. He joined Google as a software engineer in 2011 and spent over a decade there, making him one of the eight authors who hold no PhD and instead mastered his craft through pure engineering intuition.
According to the author contributions in the paper, he experimented with numerous new model variants and was responsible for the initial codebase, inference efficiency optimizations, and visualization work.
He later recalled that pivotal moment: “We were just beginning to try cutting out certain parts of the model to see how much worse it would perform. To our surprise, it actually got better.” This was the first time the hypothesis that “the recurrent structure was redundant” was validated.
In 2023, Jones co-founded Sakana AI in Tokyo with David Ha, also a former Google employee. “Sakana” means “fish” in Japanese. Ha serves as CEO, Jones as CTO, and the company’s other co-founder, Ren Ito, as COO.

Jones is now based in Tokyo and identifies on social media as a "Welsh AI researcher living in Tokyo." The company’s research approach takes a distinctly counter-trend stance: rather than relentlessly scaling compute and parameters, it draws inspiration from natural evolution, enabling a group of smaller models to collaborate like a school of fish. Notable research outputs include the Continuous Thought Machine and the "AI Scientist" project, capable of conducting end-to-end research autonomously. Recently, the company released the state-of-the-art Sakana Fugu model.

Sakana AI has raised a total of $379 million in funding, including a Series B round completed in March 2026, with Mitsubishi Electric among its investors. In March 2026, the company also secured a multi-year partnership agreement with Mitsubishi UFJ Financial Group (MUFG), which plans to use Sakana’s technology to transform its banking systems. According to reports, this partnership could enable the company, valued at approximately $1.5 billion, to achieve profitability within one year.
Jones has expressed skepticism about pure "scaling" on multiple occasions. In March 2026, at an internal banking event, he said that current AI research faces an awkward reality: despite massive inflows of investment and talent—which theoretically should drive more breakthroughs—the actual outcome may be the opposite. Investors are pushing for results, competition is driving a race to be first, and the space for researchers to engage in "free exploration" has been compressed. He noted that Sakana has always maintained a small portion of research freedom without KPIs, because the next breakthrough will inevitably come from such high-risk, long-term investments—the very approach that gave rise to the Transformer in that original Google Brain office.
He also made a frequently quoted statement: for a new architecture to truly replace Transformer, it’s not enough to be merely “better”—it must be clearly and undeniably better.
Aidan N. Gomez

Gomez was the youngest of the eight authors. At the time the paper was published, he was a 20-year-old undergraduate intern at Google Brain, pursuing a dual degree in computer science and mathematics at the University of Toronto.
According to the paper’s author contributions, he and Łukasz Kaiser spent countless nights and days building the various modules of the Tensor2Tensor framework, replacing the earlier codebase and significantly improving experimental results and research efficiency. “I just wanted to understand how the attention mechanism actually worked,” he later recalled. “I never imagined it would become the architecture for everything.” After the paper, he went to Oxford University to pursue his Ph.D., paused his studies to start a company, and only officially received his doctorate in 2024—he essentially earned his degree while building his business.
In 2019, Gomez co-founded Cohere with Ivan Zhang and Nick Frosst, positioning the company as an enterprise-grade AI service provider that deliberately avoided the costly race in consumer chatbots, instead emphasizing data privacy, on-premises deployment, and multilingual capabilities—targeting large enterprises and governments worldwide. In 2023, Gomez was named one of Time Magazine’s 100 Most Influential People in AI, and he and his two co-founders were ranked number one on Maclean’s magazine’s list of AI Trendsetters that year; in April 2025, he was appointed to the board of directors of the electric vehicle company Rivian.

This relatively unglamorous approach has instead delivered strong financial results for the company: as of mid-2026, Cohere’s annualized recurring revenue exceeded $200 million, growing sixfold over the past year, with a gross margin of approximately 70% and cumulative funding nearing $1.7 billion, valuing the company at around $7 billion. In August 2025, the company hired Francois Chadwick, who played a key role in Uber’s IPO, as its first CFO. A window for employees to sell shares on the secondary market has already opened once, and Gomez has repeatedly indicated that an IPO is “close.” However, as of now, the company has not yet filed a prospectus with regulators.
Over the past few years, Gomez has increasingly resembled a geopolitical spokesperson for AI. This week, in an article for Fortune magazine, he called on nations to confront the issue of "digital sovereignty." The article directly referenced the recent restriction of access to Anthropic’s models, warning countries not to "rent" their future to a handful of centralized tech giants, and proposed building a truly diverse ecosystem in which nations can rely on different AI providers while preserving their own values, languages, and legal systems.
He has also publicly stated that concerns about existential risks from an “AI doomsday” are exaggerated, and that his greater concern is the real risk of misinformation being amplified automatically on social media. Gomez now speaks not just about the models themselves, but about who gets to decide what kind of AI the world uses.
Łukasz Kaiser

Kaiser is Polish and initially trained in theoretical computer science, including logic, automata theory, algorithmic model theory, and game theory. He earned dual master’s degrees in mathematics and computer science from the University of Wrocław, completed his Ph.D. at RWTH Aachen University in Germany, and later held a tenured position at the French National Center for Scientific Research (CNRS) and Paris Diderot University, focusing on pure theoretical research in logic and automata theory. He later shifted toward applied work, spending nearly eight years at Google Brain, where he was a co-author of TensorFlow and co-authored early papers with Samy Bengio on “Can Active Memory Replace Attention?” and with Ilya Sutskever on “Neural GPU Learning Algorithms.”

According to the author contributions in the paper, he and Aidan N. Gomez spent countless nights and days building the tensor2tensor framework, significantly improving experimental results and research efficiency.
Among the eight authors, he was the only one who never started a company and remained in a large laboratory conducting pure research.
In 2021, he joined OpenAI, before ChatGPT had been released. At OpenAI, he contributed to the development of Codex (which later became the technical foundation for GitHub Copilot) and the accompanying HumanEval programming benchmark, as well as research on the GSM8K math problem dataset. This work early demonstrated that "allowing the model to reason longer and sample more times" significantly improves accuracy—laying the groundwork for what would later become the reasoning model paradigm.

He was also one of the named authors of the GPT-4 technical report and later became a core contributor to OpenAI’s first reasoning model, o1 (released in September 2024), in a role regarded as “lead researcher,” a trajectory that continued through o3 and subsequent reasoning paradigms, leading up to today’s GPT-5 series.
He recently discussed on Matt Turck’s MAD Podcast that transformers have been mathematically proven to solve any problem, provided the model is allowed to generate a sufficient number of intermediate reasoning steps. In a sense, this is a belated, more precise commentary on a paper from nine years ago.
Illia Polosukhin

Polosukhin is from Kharkiv, Ukraine, where he studied applied mathematics as an undergraduate and was a champion of the International Collegiate Programming Contest (ICPC). He recalls that after watching The Matrix at age ten, he developed an almost obsessive interest in artificial intelligence. In 2014, he joined Google, where he worked on research related to TensorFlow and conducted studies in machine reading comprehension and question-answering systems.
According to the author contribution statement, he designed and implemented the original Transformer model together with Ashish Vaswani, with his primary responsibility being to validate the effectiveness of this architecture on machine translation tasks.
After publishing the paper, he left Google in 2017 and co-founded an artificial intelligence company initially called NEAR.AI with Alexander Skidanov. But they soon realized that building decentralized infrastructure might be more interesting than developing models, so the company transitioned into the blockchain project NEAR Protocol around 2018.

NEAR employs a sharding technology called Nightshade and offers an Ethereum-compatible Layer 2 network through Aurora. The mainnet launched in 2020 and has raised over $530 million from institutions including a16z, Coinbase, Tiger Global, Hashed, and Dragonfly Capital.
Today, Polosukhin is trying to reconnect his two original identities: In March 2026, he told the media, “The future users of blockchain will be AI agents, not humans,” positioning NEAR as the “settlement layer” for the agent economy. In April of the same year, he publicly called for a more robust regulatory framework to address autonomous AI agents; he argued that existing institutions and systems are not yet prepared to handle accountability and systemic risks posed by such systems, urging the creation of clearer accountability mechanisms and “human-in-the-loop” oversight.
He is currently based in Portugal. Worldwide, he may be the only person who has both authored foundational LLM papers and run a blockchain company valued at billions of dollars.
Eight paths, continue exploring
In March 2024, at the NVIDIA GTC conference, seven of the eight authors (Niki Parmar was absent due to unforeseen circumstances) appeared together for the first time as a group in an interview with Jensen Huang.

Jensen Huang said: "Everything we enjoy today can be traced back to that moment."
At the end of the conversation, he presented each of them with a signed NVIDIA DGX-1 supercomputer commemorative plaque engraved with “You transformed the world.” In November of the same year, Japan’s NEC C&C Foundation awarded the annual C&C Prize to the eight-member “Transformer team,” who shared the stage with three senior engineers who had pioneered transoceanic submarine fiber-optic cable transmission technology. Two groups of infrastructure builders from entirely different fields were honored together under the same award.

Nine years later, these eight life paths have diverged to places that rarely, if ever, intersect: enterprise services in Silicon Valley, evolutionary algorithms labs in Tokyo, molecular biology companies in Berlin, blockchain protocols in Portugal, and several leading AI labs still reconfiguring themselves this week.
But if you look at everything they’ve said over the years, a recurring judgment emerges: no one truly believed that Transformers would be the end point.
Aidan N. Gomez said the world needs something better than Transformers; Llion Jones said the next architecture must be “clearly and unquestionably better” to replace it; Łukasz Kaiser is still using mathematical language to clarify how far this nine-year-old architecture can still take humanity.
Perhaps this is the most enduring legacy of the paper: its eight authors, scattered across the globe, none of whom have stopped seeking the next answer.
Reference link
https://www.wired.com/story/eight-google-employees-invented-modern-ai-transformers-paper/
https://x.com/TylerMaran/status/2067772926695522454
https://www.nvidia.com/en-us/on-demand/session/gtc24-s63046/
This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), authored by someone interested in AI.
