In recent years, Google's Gemini model has underperformed, falling from once-revered frontier AI to a subject of online mockery as the "North American soybean pack." The model frequently makes basic errors such as faulty mathematical reasoning and flawed image generation, ranks poorly in programming benchmarks, and lags behind GPT and Claude in knowledge base updates. The root cause lies in Google’s disorganized internal structure, where different departments vie to make AI their own gateway, leading to fragmented resources and uneven compute allocation. Key researchers have left en masse, including the authors of Transformer and the Nobel laureate behind AlphaFold, all poached by competitors. Although Google has released Gemini 3.6 Flash and is already training Gemini 4, returning to the pinnacle of AI remains a formidable challenge.Author and source: Chaping X.PIN
I admit that Gemini was strong last year, but now its reputation has become like a North American soybean package...
I admit that Gemini was strong last year, but now its reputation has become like a North American soybean package.

I asked it math questions, and suddenly it spat out poop.

It asked it to draw a photo of a beach, but the user immediately assumed it was a pedophile.

Because of these vague statements, Gemini was directly labeled as a "North American soybean package" online.

In addition to the model itself taking a sharp downturn, Google has recently seen frequent changes among its research staff.
Two years ago, one of the authors of Transformer, who was recruited for $2.7 billion, has now left and joined OpenAI.

Even the Nobel laureate who worked diligently for nine years at AlphaFold has been sending subtle signals to Anthropic.

Everyone had been eagerly awaiting the release of Gemini 3.5 Pro to redeem themselves, but from May all the way to July, it still hadn’t been released.
A few days ago, there was finally some movement, but what was delivered wasn't the flagship Pro model—it was a Flash model focused on affordability.

What is the capability level of this model?
To put it this way, while other models are benchmarking against Fable and GPT-5.6 Sol, Gemini 3.5 Flash is comparing itself to GPT-5.6 Luna and Claude Sonnet 5.

Just from this comparison, we can see Google's expectations for this new model...
In the areas of AI programming and agents, which are currently the top priority for everyone, Gemini has fallen far behind.
On the top pages of major rankings, you might not even find its name.


In real life, no one would really use Gemini to write code, would they? I previously had it handle a project migration, but it just focused on providing emotional support—loading me up with compliments while getting almost nothing done.

Even Google's most prized global knowledge capability swings wildly between夯 and 拉.
It currently lags in breadth; previously, DeepSeek openly acknowledged in its paper that it falls short of Gemini in terms of knowledge breadth.
DeepSeek-V4-Pro-Max significantly outperforms all open-source models on SimpleQA-Verified but still lags behind the leading proprietary model, Gemini-3.1-Pro.

When you ask it obscure, old-school facts, only Gemini might remember.
However, once you discuss things that have emerged in the past two years, Gemini's responses become extremely odd.
For example, while writing this article, we had Gemini help us organize the outline.
The result was that, with no contextual relevance, Gemini suddenly brought up Claude 3.5.

Yes, by default, Gemini only recognizes Claude up to version 3.5.
This is because, in the current Gemini models, whether it's the 3.5 Flash released in May this year or the Gemini 2.5 Pro released in June last year.
The built-in knowledge bases of these models only contain information up to January 2025.

So when Claude 3.7 was released in February 2025, it didn’t recognize it at all—only knew about 3.5, which was released in 2024—and thus confidently started making things up.

This means that when you normally chat with Gemini, if the conversation touches on events from the past year and a half, Gemini may become confused or erratic.
If you want it to learn the latest information, you must enable web search—but often, after enabling web search, the model’s performance declines because it absorbs too much information, becoming stuck between two extremes.

Meanwhile, the other two major models, GPT and Claude, updated their knowledge bases to the end of last year long ago—only Google continued using its outdated knowledge base until the recent release of Gemini 3.6 Flash, which finally updated its knowledge to March 2026.

It’s hard to imagine that Google, which started with search, is being outmaneuvered in the very area of information it excels at.
So, what exactly happened with Gemini?
After scouring reports from various authoritative media outlets, we find that there is nothing new under the sun—Gemini’s recent underperformance is, in fact, a reflection of the underlying disorganization within Google’s AI organizational structure.
Unlike purely large model companies like OpenAI and Anthropic, for Google, Gemini has never been just a chatbot.
It is both a cutting-edge model in DeepMind’s hands, a weapon for Google Search to maintain its search gateway, an API sold by Google Cloud to enterprise customers, and is also set to integrate into Android, Chrome, Workspace, Pixel, and Google Home.

Each business line wants to make AI its next-generation entry point and believes its own use cases are the most important.
So who is the most important one?
The Information previously reported that colleagues at Google AI Labs once developed an AI note-taking app called NotebookLM, which upset colleagues working on Gmail and Google Docs upon its release.
Who told you to do this? Don’t you understand workplace etiquette? Aren’t you taking away your own work?
There were even rumors that Workspace employees had considered shutting down NotebookLM.

In addition, the Google Cloud team and the DeepMind team also had a clash.

The Pixel phone team wanted to add AI features to Pixel phones, but was internally instructed not to compete with the Gemini assistant.

After the company grew too large, each team began protecting its own territory, and every group worried that others’ products would encroach on their position.
AI hasn't truly begun to change the world yet, but inside Google, the race for the steering wheel has already started.
The English philosopher Hobbes described in Leviathan a state of war of all against all.

In short: you want everything, but there isn't enough cake to go around, so you're stuck in endless squabbles.
Today, Google is like this—the boundaries between departments have become as impenetrable as the Chu River and Han Boundary, and everyone tightly guards their own limited share of resources.
This trend has even affected Google's internal allocation of computing power.
Logically, Google should be the last company to lack computing power.
It has its own TPUs, cloud infrastructure, and data centers, and also sells TPUs to external companies like Anthropic, Meta, and Midjourney.
But how to balance business and research became a major issue within Google.
Inside Google, some AI researchers feel that the company prefers to allocate computing power to paying customers, who are truly the priority.

These internal complexities have also led to a number of employees leaving.
Even many AI researchers have found that, rather than going through the lengthy process of applying for computing power and waiting for approval at Google, it’s simpler to just quit outright and work for themselves.
Starting your own business and securing computing power yourself is much faster than waiting around at Google.

If computational power is Google’s limitation for ordinary researchers, for many top-tier researchers, Google may also fall short when it comes to financial resources.
As a supergiant with a market capitalization exceeding $4 trillion, Google compensates these top talents primarily with restricted stock tightly tied to the company, in addition to cash.

Can this stock make money? Yes, it can make money. Can it make a lot of money? That’s a bit difficult.
Even if you work as hard as you can, could Google's market capitalization possibly multiply several times in the short term?
But OpenAI and Anthropic are different.
Both have been quietly preparing for their IPOs lately—if you act now, you might get a chance to secure some early shares, and once the listing is complete, you could significantly boost your wealth, which is far more appealing than working a regular job.

Of course, Google is not unaware of these issues.
Recently, we restructured our organization by creating a dedicated "Mid-Training" team between the standard Pre-Training and Mid-Training phases to specifically enhance the model's programming capabilities.

Although there is much discussion about Gemini 3.6 Flash, Google has now confidently stated that it has begun preparing to train Gemini 4.

Whether Google can reclaim its former glory may depend on this round.
Although we’ve discussed many issues with Gemini today, you know I’ve always been a fan of Google.
I really loved the 2.5 Pro from last year, and even now, I still feel that the text written by Gemini has a more human touch.
This sense of being alive becomes even more precious in an era when everyone is competing to master coding skills.
