Google DeepMind has released three Gemini model updates: the flagship model, Gemini 3.6 Flash, reduces token usage by up to 65% in coding tasks, with pricing at $1.50 per million tokens for input and $7.50 for output—making it more cost-effective—and demonstrates significantly improved performance on benchmarks such as DeepSWE. The lightweight variant, Gemini 3.5 Flash-Lite, achieves output speeds of up to 350 tokens per second, with pricing as low as $0.30 per million tokens for input and $2.50 for output, outperforming the larger Gemini 3 Flash across multiple tests. The security-focused model, Gemini 3.5 Flash Cyber, specializes in vulnerability remediation. Google has also announced its most aggressive pre-training initiative yet, targeting Gemini 4. For users, the key benefits are comprehensive reductions in cost, token consumption, and Agent operation expenses.Author and source: AI New Era
For those of us who have spent real money to buy tokens and build agents, today’s launch really comes down to just one word: reduction.
Three in a row—Google is really not holding back anymore.
Today, Google DeepMind unleashed a series of major announcements, revealing three powerful innovations at once:
- Gemini 3.6 Flash
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash Cyber

A stronger主力, a faster lightweight version, and a security specialist focused on vulnerability patching.
The three Gemini models point to the same goal: making AI agents running in production faster, smarter, and more cost-effective.

On the same day it was released, Google also dropped an even bigger bomb—
Google DeepMind's most aggressive pre-training yet has begun, targeting the next-generation model, Gemini 4.

On one side, Gemini is launching three consecutive updates; on the other, Gemini 4 has just begun.
Google's message could not be clearer: this ultimate poker game of AI has only just begun.
Google unleashes a late-night surprise, launching three Gemini models simultaneously.
Three Gemini products, to be broken down one by one.
Gemini 3.6 Flash, saves up to 65% in tokens
Let’s start with the real star of this show—Gemini 3.6 Flash.
This time, it highlights a key feature: massive Token savings.
According to the Artificial Analysis Index, it uses 17% fewer output tokens compared to the previous generation, 3.5 Flash.
On coding benchmarks like DeepSWE, up to 65% can be saved.

It not only speaks less but also requires fewer reasoning steps and tool calls to complete multi-step tasks, taking a shorter path—making it both faster and more efficient.
Moreover, Gemini 3.6 Flash is not only more token-efficient but also cheaper:
Input $1.5, output $7.5 per million tokens—cheaper than 3.5 Flash.
More savings, lower cost, and even better performance—that’s real诚意. The 3.6 Flash outperformed its predecessor on several key tests—
- DeepSWE: Programming Test 37% → 49%
- MLE Bench: Machine Learning Research Test 49.7% → 63.9%
- OSWorld-Verified: Enables models to directly operate a computer, improving test scores from 78.4% to 83%.
- GDPVal-AA v2: Completed 1,421 knowledge-based tasks, over 70 more than the previous generation.

In the demo below, 3.6 Flash demonstrates exceptional multi-agent orchestration, effortlessly handling complex code migration tasks while outperforming 3.5 Flash in both response speed and code quality.

With Gemini Canvas, 3.6 Flash has streamlined the 3D workflow, enabling users to instantly create a professional, photorealistic texture extraction tool.
It can also transform into an "AI Interaction Designer," effortlessly creating immersive themed workspaces.


3.5 Flash-Lite, outperforming the giant
The second model, Gemini 3.5 Flash-Lite, takes a different approach: it’s designed to be fast and affordable.
3.5 Flash-Lite achieves an output speed of 350 tokens per second, making it the fastest in the 3.5 series.
The prices are also ridiculously low: $0.30 per million tokens input and $2.50 per million tokens output.
Fast and affordable, it’s naturally designed for high-frequency, high-volume use cases like bulk document processing and Agent search.
Most impressively, it actually turned around and stepped on its own "big brother."
In multiple programming and agent tests, this lightweight small model outperformed the much larger Gemini 3 Flash—
- SWE-Bench Pro: 54.2% vs. 49.6%
- OSWorld-Verified: 74.0% vs 65.1%

3.6 Use Flash as the "brain" to decompose tasks, and Flash-Lite as "avatars" to handle them in bulk.
One handles thinking, many handle running—easily smoothing out high-concurrency pressure.
In the official demo, Flash-Lite seamlessly combines two approaches to deliver 25 high-availability website designs in one go—so fast it’s almost like generating a batch with just a thought.
Flash-Lite is now live on the Gemini app and will gradually be available on Google Search.
Flash Cyber: Can fix vulnerabilities, but not everyone can use it
Gemini 3.5 Flash Cyber, a powerful model specialized in cybersecurity.
Its only task is to find and fix vulnerabilities.
Right now, a troubling reality is that AI is finding vulnerabilities faster than existing systems can patch them.

Google integrated this model into the code security agent CodeMender, leveraging multiple Cyber agents to collaborate and achieve a new SOTA on the CyberGym security benchmark—at a lower cost than larger models.
Use a lighter model to handle the tasks that require the highest precision.
However, this model is not accessible to the average person for now.
Gemini 4 is live—the most aggressive pre-training ever
Although Gemini 3.5 Pro hasn't been acquired yet, training for Gemini 4 has already begun.
This triple punch is just the appetizer, but the final official announcement is the real highlight.
Google has confirmed that it has internally launched its most aggressive pre-training effort yet, with the goal of reaching Gemini 4.

The expert said that, following Google's usual six-month training cycle, we might see Gemini 4 by the end of the year.

For those of us who have spent real money to buy tokens and build agents, today’s launch really comes down to just one word: reduction.
The price has dropped, Token consumption has decreased, and the cost to run an Agent has gone down.
As for the long-awaited flagship and the newly launched Gemini 4, both are still "futures."
What’s practical and affordable is what you can put in your wallet today. Even the flagship product must first come down to earth.
