Truly impressive!
This time, Google has finally laid its cards on the table.
In the past couple of days, a new model named "gemini-3.8-flash" has quietly entered the Large Model Arena.
After several rounds of hands-on testing against AI experts and developers, they nearly gasped in surprise—
This is almost certainly Google DeepMind’s next flagship Gemini 4 Pro, after six months of intense development!

A widely circulated benchmark image that is simply terrifying.
Gemini 4 Pro is completely dominating, outperforming GPT-6 Astra and Claude Fable 5.1 across coding, agents, reasoning, and other areas.

After trying it out, developer Vidhi was completely amazed by Gemini 4 Pro.
An SVG image of a cat proved it all. GPT-6 Astra directly generated a "big cat" (cat-like but more like laofu).


As rumored, Google internally implemented RSI to enable Gemini 4 Pro to directly compete with Astra/Fable.
The long-dormant giant has made a powerful return!
Gemini 4 Pro anonymously outperforms Astra and Fable
The AI community hasn't been this excited in a long time.
Keep in mind that Google's major update to Gemini 3.1 Pro was already seven months ago (February 19).
At Google I/O, Pi Chai previously announced that Gemini 3.5 Pro would launch in June.
Unexpectedly, this major announcement was delayed again and again, ultimately being canceled!
At the beginning of this month, the WSJ exclusively reported that Gemini 3.5 Pro was discontinued, for a simple reason—
The level of evolution of 3.5 Pro doesn't even compare to Flash.
Fortunately, Gemini 4 performed well during pre-training, and post-training is progressing smoothly.
Now, Gemini 4 Pro has appeared on Arena dressed in the "clothing" of Gemini-3.8-Flash, making its debut at the first checkpoint.
After all, the Gemini 3.8 Flash model was released in early September; its reappearance under the same name is highly unusual.
A large group of developers who have tried it out after hands-on experience were impressed by the 4 Pro, even feeling it surpassed the Astra and Fable.

From the leaked benchmark charts, it is evident that the 4 Pro delivers comprehensive dominance, with an outstanding performance record—
DeepSWE v1.1: On AI agent coding tasks, the 4 Pro achieved a maximum score of 88.%, nearly 2% higher than Astra;
GDPval-AA v2: Among real knowledge tasks, only 4 Pro achieved an Elo of 2064;
Terminal-Bench 2.1: The 4 Pro has the strongest terminal coding capability at 95.3%;
OSWorld-2.0: In terms of computer usage capabilities, the 4 Pro achieved 86.8%, outperforming Astra and Fable.
If the scores on this table are accurate, Gemini 4 Pro is truly the strongest AI on Earth.
In terms of pricing, it is also the most cost-effective among the "Big Three" compared to Astra and Fable, at $2.25 per million tokens input and $11.25 per million tokens output.

AI researcher Qwinah discovered after accessing Gemini 4 Pro's backend that it has a 10-million-token input limit, a 256,000-token output limit, permanent cross-session memory, and can connect to the internet without requiring an API.

First-ever network test of the 4 Pro, trending on热搜
What’s truly more impressive than score-running is a sweeping, real-world test across the entire network.

UI web design with significantly enhanced sophistication
After testing it, developer Bee exclaimed, "Gemini 4 Pro is just unreal..."
In just 14 minutes, 4 Pro created a creative showcase webpage themed around sketches, pencils, and graphite textures.
It brilliantly conceptualizes the webpage's "scroll as stroke" idea, with lines gradually darkening as you scroll down, offering an incredibly smooth interactive experience.

This is a work website that combines cyberpunk and retro-futurism.
During the design process, the 4 Pro integrated a 3D grid, data dashboard, and interactive 3D models on the home screen.
3D+SVG, absolutely smashing
A classic "pelican riding a bicycle" test that's gone viral across the internet—Gemini 4 Pro’s output is incredibly impressive.
Color scheme, day/night switching, headlights, anatomical annotations, cadence control—the overall completion is on par with top-tier large models.
After experiencing it, Pankaj Kumar summarized several changes:
Very fast, with noticeable improvements in SVG and 3D generation; accurately understands complex requests with a single prompt.
In fact, it can now directly generate small games with complete interactive logic and appealing visual effects.


Let’s take another look at the pixel-style 3D pagoda test—it now feels truly immersive.

Below is the 3D model of the Airbus H145 helicopter created in 10 minutes using the 4 Pro.

Additionally, in a 3D flight simulation test, the 4 Pro's visual performance significantly surpassed that of the Gemini 3.8 Flash.
From this, it is clear that Gemini 3.8 Flash and Gemini-3.8-Flash on the leaderboard are not the same model.


One-click playable games
Gemini 4 Pro also excels in game generation.
"My World" matches the quality of Astra's previous test release in every aspect, from page construction to the design of each module within the game.
There's also this 3D go-kart racing game, which looks more modern than once-popular racing games like KartRider; and this was created by Gemini 4 Pro in a very short time.
What Google is truly betting on is called RSI.
Once this RSI cycle truly gains momentum, the pace of model upgrades could change completely.
Last month, Jasjeet Sekhon, Chief Strategy Officer at Google DeepMind, laid it out plainly at an event in Berkeley—
RSI (Recursive Self-Improvement) has become a critical component in the logic behind massive AI investments.
The entire industry is betting on the continued upward trajectory of the capability curve.
And if true recursive self-improvement is achieved, this curve could accelerate even further beyond exponential growth.
Recently, more radical claims have been circulating online:
Gemini 4 completed its pre-training ahead of schedule because Google DeepMind achieved an "RSI feedback loop" during training.

Interestingly, on the 14th, Google also released a study titled Dream-RSI.
It specifically investigates how an agent can continuously improve its search strategy during exploration, upgrading its approach for the next round based on experience.
It is confirmed that Google DeepMind has publicly placed significant emphasis on the RSI.
Moreover, there is increasing evidence of AI participating in the improvement of AI.
Are the "Big Three" coming back?
So, Gemini 4 Pro is facing a completely transformed table.
OpenAI has Astra; Anthropic has Fable 5.1.
The two giants have escalated competition beyond simple chat to include long tasks, agents, code, reasoning, and even autonomous scientific research.

If Google merely releases a model that is "stronger than Gemini 3.8," it is no longer enough.
Gemini 4 Pro must reprove one thing: Gemini is still at the forefront!
For Google to maintain its current market capitalization of $4.23 trillion, staying at the forefront of AI capabilities is crucial.
Reference materials:
https://x.com/HarshithLucky3/status/2100524699780600136
This article is from the WeChat public account "New Intelligence Yuan" (ID: AI_era), authored by Peach and Marco.
