Just now, Google has made a comeback!
Tonight, Google officially released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, specifically designed for cybersecurity.
This is the third Flash version released by Google within just six weeks.
The previous two updates seemed like mere warm-ups, because this time, Google has directly pushed cost-effectiveness to the forefront—
Single task cost: $0.58! Input just $0.75 per million tokens!
This tiny model, priced at a fraction of the cost, has nonetheless achieved performance levels on key benchmark tests that rival those of the world’s top flagship models: Opus 5, GPT-5.6 Sol, and Grok 4.6.
Google DeepMind expert Yao Shunyu commented: For the model, this is just a small step; but for RSI, it’s a giant leap.

On X, developers have exploded.
After real-world testing, someone asked the burning question: “Oh my goodness, did Google ship the wrong product? This is clearly Gemini 3.5 Pro, never officially released—paying Flash prices for Pro-level performance!”


In the DeepMind team, someone probably hasn't slept since July.
While everyone is still processing Fable 5.1, Google has once again turned the table.
Google's "King of Efficiency" is back: The King of Value
After a long period of silence, Google has announced to the world—using an intensely competitive approach—that the big boss is back.

To understand the impact of Gemini 3.8 Flash, we must first examine the current landscape of the AI market.
Originally, in the Artificial Analysis test, an extremely strict barrier had already been established. To achieve top-tier intelligence (an Intelligence Index of around 60), you must incur a high Token cost.
As a result, Gemini 3.8 Flash is like an assassin—completely ruthless.
Under Artificial Analysis's high-reasoning mode, Gemini 3.8 Flash's IQ surged to 59!

What does this mean? This is just one step away from the top-tier GPT-5.6 Sol and Grok 4.6, and even surpasses Claude Opus 5 in certain dimensions!
And the cost of achieving all of this? Just $0.58 per task.
Specifically, the input cost remains at $0.75 per million tokens, and the output cost is $3.75 per million tokens.
Google created a chart: On the Pareto frontier of intelligence versus cost, the blue dot representing Gemini 3.5 Flash sits at the top-right corner of the DeepSWE software engineering benchmark, leaving the second-place model far behind.

Gemini 3.8 Flash is not only intelligent but also incredibly fast, generating text at an astonishing rate of 305 tokens per second, while the next-tier model barely reaches 154 tokens per second.
Fast, smart, and cheap as if it were free—that’s the kind of model humans can truly use every day.

In particular, on the long-cycle software engineering benchmark (DeepSWE v1.1), 3.8 Flash achieved an impressive score of 73.7%.
In this test requiring AI to autonomously solve complex engineering problems and write end-to-end code, it not only outperformed most expensive state-of-the-art large models but also cost only a fraction of theirs.
Keep in mind, this is only the "Flash" version, not the "Pro," let alone the "Ultra."
This time, Google once again let a small model steal the spotlight from its flagship model.
Someone personally tested Gemini 3.8 Flash and Opus 5 on the same task.

This significant leap in programming ability is no accident.

This significant leap in programming ability is no accident.
According to reports, during internal testing of Google’s code tool Jetski, Google engineers have even preferred using 3.8 Flash over Anthropic’s Opus model.

Behind this is Google’s heavy investment in reinforcement learning since the beginning of the year—the final refinement phase of model training, where models learn to perform tasks through trial and error.
Google DeepMind has formed a code strike team focused on enhancing the capabilities of programming models, led by DeepMind’s CTO and directly overseen by co-founder Sergey Brin.
Their goal is to force recursive self-evolution, forcibly transforming the programming model into an fully automated AI researcher, completely closing the entire R&D loop.

In the internal memo, Google directly stated:
To win the final sprint, we must urgently close the gap in agent execution and turn our models into primary developers.

Additionally, Google has recruited Barret Zoph, co-founder of Thinking Machines Lab and former head of post-training at OpenAI, who now serves as Vice President of Research, overseeing reinforcement learning and post-training initiatives. This move clearly signals their determination to see this through to the end.
Last month, co-founder and Nobel laureate Demis Hassabis stepped down as CEO, and his successor, Koray Kavukcuoglu, immediately declared: speed up, speed up, and speed up again.
Is it baseless inflation, or genuine strength?
However, amid the praise, Google was widely criticized for celebrating too soon.

The most frustrating is Meta—
Meta's Muse Spark 1.3 and Gemini 3.8 Flash go head-to-head, with the former scoring 62 on Artificial Analysis's intelligence index, tying with Claude Fable 5 and entering the top tier.
Muse’s input cost is 8 times lower, and its output cost is nearly 12 times lower than Fable 5, delivering unmatched value for money.
Meta’s Chief AI Officer, Alexandr Wang, directly mocked Gemini: on the Artificial Analysis intelligence index, Gemini is nothing but a footnote, forced to breathe in the exhaust fumes of other models.


Muse Spark 1.3 scores higher than GPT-5.6 Sol, Grok 4.6, and Gemini 3.8 Flash, but is currently only available in a limited preview.
Some have pointed out that while the benchmark results for 3.8 Flash look impressive, real-world performance may not be as strong.
Indeed, Gemini 3.8 Flash outperformed GPT-5.6 Sol and Opus 5 in tests such as Terminal-Bench 2.1 and HLE, but the significant score gap between TBench 2.1 and TBench 4 reveals the potential presence of "benchmark manipulation."
In other words, when faced with unfamiliar, real-world zero-shot challenges not included in the training set, 3.8 Flash is still likely outperformed by parameter-heavy models like Opus 5.
But Google's solution is extremely clever—diligence compensates for lack of talent.
As stated in Google's official blog: "These performance improvements stem from core design choices: 3.8 Flash works harder."
When faced with complex tasks, 3.8 Flash is designed to perform additional reasoning steps and iteratively invoke tools. When encountering unfamiliar questions, it does not give up; instead, it consumes more tokens to conduct rapid self-verification and reflection internally.
Even though it consumes more tokens through multiple rounds of reasoning, its base price is so low ($0.75 per million tokens) that, overall, it’s still significantly cheaper than calling GPT-5.6 directly!
This strategy directly breaks the past single-dimensional competition of "larger parameters = stronger capability."
It demonstrates that, in a perfect agent loop, speed and extremely low inference costs are themselves a form of powerful intelligence.
Why send a Flash? The "canceled" Pro version
This time, did Google really not ship the wrong item?
Gemini 3.5 Pro hasn’t even appeared yet. The next-generation flagship, Gemini 4, has impressive pre-training data, but its post-training phase isn’t complete yet.

In fact, Google's prolonged delay in releasing the Pro version has a bittersweet backstory.
Pi Chai boasted in May this year that the more powerful Pro series would be launched "next month," but has continually delayed it.
Actually, Google internally had several candidate models for 3.5 Pro, but all were ultimately discarded—because none demonstrated a significant enough advantage over the current Flash series!
Meanwhile, the highly anticipated next-generation flagship, Gemini 4, although performing well in pre-training evaluations, is currently stuck in the most critical post-training phase.
In summary, a series of unexpected events led to the rapid iteration of the Flash series.
Found a vulnerability in 2 hours that would have taken months to discover: dominating the cybersecurity leaderboard
Is Google just driving down prices for 3.8 Flash on routine tasks this time?
Far beyond that.
The real highlight of this launch is its twin brother—Gemini 3.8 Flash Cyber.
Even developers are amazed: “What kind of security task could possibly warrant Google creating a dedicated Cyber variant for Flash?”
Google's answer is: to combat future AI hackers.
On the CyberGym benchmark where the vulnerability was discovered, 3.8 Flash Cyber scored 86.2%, topping the leaderboard and far outperforming previous large models.
Moreover, real-world code is not limited to C/C++.
In a comprehensive test across 20 programming languages within Google’s internal codebase, this model achieved a success rate of over 70% in identifying widespread vulnerabilities.
Even more impressive are its real-world battle results.
The Chrome security team tested it and found that it generated 2.6 times as many correct patch fixes as the best previously available commercial models, which were much larger.
Wiz security testing found that its recall rate improved by 7.5–9.7% during penetration testing, while costs decreased by a factor of 2.3 to 5.2.
Most surprisingly, Google Cloud’s vulnerability research team used it to discover a critical foundational vulnerability in just two hours—a task that typically takes top human experts months to accomplish.
On CWE-Bench (patching capability test), 3.8 Flash Cyber achieved a first-pass rate of 47.2%, nearly matching the leading state-of-the-art large models (47.8%), while costing only a fraction of the price.

Precisely because of the extraordinary offensive and defensive capabilities of 3.8 Flash Cyber, Google has chosen not to fully open-source it, but instead to make it available only to trusted organizations through the new Fairwind program.
OpenAI, Anthropic, and Google: The Big Model Triad Reaches a Turning Point
The release of Gemini 3.8 Flash reveals that today's three AI giants have embarked on three distinctly different evolutionary paths.
Anthropic (Claude series) is relentlessly focused on knowledge reliability and stability.
In tests emphasizing knowledge coverage and factual accuracy, such as AA-Omniscience, Claude still delivers a decisive advantage.
The top seven are all Claude models, which are now clearly focusing on enhancing their error-free performance in enterprise tasks, aiming for consistent and reliable outputs.
OpenAI (GPT series) is betting on "deep reasoning and insight."
In the CritPt test favoring physical and mathematical derivations, GPT-5.6 Sol leads by a wide margin.
OpenAI appears to be pursuing a pure emergence of intelligence, requiring models to independently "figure things out" like scientists, even without being told the answers.
Google (Gemini series) creates an "endless, hyper-fast worker loop."
This time, Google offered a third answer: Maybe I can't solve the hardest physics problems all at once, but I react extremely quickly and at an extremely low cost.
In the age of agents, I can think 100 times per second—write code, run it, encounter errors, fix them—using brute-force iteration at an extremely low cost to forcefully arrive at the correct result.
Agen encounters practical challenges that require repeated trial and error, tool invocation, and dynamic adjustment.
Gemini 3.8 Flash, on the other hand, is perfectly tailored for this kind of "long-range workflow."
In Google Antigravity, you can generate a complete game with puzzles, environment textures, and 3D levels using just a single prompt plus a loop command with 3.8 Flash.
You can use it to instantly generate a DOS version of Google Maps with street view and navigation.
Clearly, Gemini 3.8 Flash's ease of use and cost have surpassed a critical threshold.
The arrival of Gemini 3.8 Flash feels like Google declaring to the entire industry: large models are no longer limited to only the biggest players—they're racing toward accessible computing power for all.
Tasks that once required tens of thousands of dollars and days of effort can now be completed in minutes for the cost of a few cups of coffee.
The era of the super individual for ordinary people has truly arrived.
This is the awakening moment for small models.
Reference materials:
https://x.com/alexandr_wang/status/2095266112212754659?s=20
https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/
Edited by: Aeneas David
This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation.
