Google to Launch Gemini 3.8 Flash Tomorrow, Focusing on Coding and Video Understanding

icon MarsBit
Share
AI summary iconSummary
Google is set to launch Gemini 3.8 Flash tomorrow, a major token launch event in the AI space. The model, codenamed 'Skimaki,' enhances coding and video understanding while reducing token usage by up to 88%. Gemini 3.5 Pro has been canceled, with development now focused on Gemini 4.0. On-chain news indicates that Google is also introducing agentic video understanding for dynamic content analysis.

Claude 5.1 has just been released, while OpenAI's Astra is right on its heels.

Now, even Google has made a new move. Latest reports reveal that Gemini 3.8 Flash will launch tomorrow.

Google

Google

It seems the AI community is about to have a big celebration.

Gemini 3.5 Pro has been discontinued; the next major version is 4.0.

Many are more curious about when Gemini 3.5 Pro will arrive, rather than the release of Gemini 3.8 Flash!

Sorry, no need to wait—Google has given up...

Google

Nearly four months have passed since the I/O conference preview in May, and Gemini 3.5 Pro has evolved less than Flash.

Google also had no choice but to directly "sacrifice the pawn."

Fortunately, Gemini 4.0 performed excellently during pre-training, and post-training is progressing smoothly.

Google

Google

This exclusive report from the WSJ also highlights Gemini 3.8 Flash (codenamed "Skimaki")'s powerful programming capabilities.

In internal benchmark tests against Google's coding tool Jetski, 3.8 Flash reportedly outperformed the Opus model by a wide margin.

Moreover, this model has been under development for months, long before the "earthquake" at Google's leadership.

Hassabis stepped down, and Chief Scientist Jeff Dean left, causing many to worry about Gemini’s future.

Google

However, it appears that everything is proceeding smoothly internally.

Currently, Google plans to launch Gemini 3.8 Flash first because this model has fewer parameters, requires less computation, and is easier to deliver results with.

According to insiders, since the beginning of this year, Google has allocated more human and computational resources specifically to enhance Gemini’s code-writing capabilities.

In particular, increased investment has been made in "reinforcement learning," enabling AI to truly acquire new skills through continuous trial and error.

Moreover, Google DeepMind and other labs are advancing multiple projects simultaneously, so if one doesn’t work out, another can take its place.

Gemini 3.5 Pro fell short of expectations, but Gemini 4.0 has not yet been released.

Today, Silicon Valley’s two leading companies, OpenAI and Anthropic, have demonstrated strong capabilities in cutting-edge models and programming agents.

According to information, OpenAI’s next-generation Astra also employs a novel reasoning method called "recurrent depth," which reduces the number of thinking tokens while significantly enhancing performance.

This ace will be revealed this week.

Google

At this moment, Google must deliver something.

After Pi's promise, for months there was nothing but the release of a 3.7 Flash model—nothing else to excite the AI community.

Perhaps Gemini 3.8 Flash will be unveiled tonight.

Google

Gemini is seeing a video for the first time.

Before this release, Google kicked off today with a preview, introducing a new capability to Gemini—

Agentic Video Understanding

This feature is a game-changer.

Upload a video from the I/O conference: for the same model, Gemini 3.7 Flash, token consumption drops by 88% directly.

Google

Google stated on its official blog that the Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models all support agentic video understanding, accessible via Google AI Studio and the Gemini Enterprise Agent Platform.

Turning on this switch on standard video analysis benchmarks can reduce token consumption by up to 88%, lower analysis costs by up to 66%, and increase accuracy by up to 7%.

Google

And there are no additional fees—it still uses the standard API token pricing. Saving money and improving accuracy often conflict, but this time, they didn’t.

Google

Because Gemini learned to "scrub the progress bar."

The previous approach was called "static" processing, where the model passively received a fixed-frame-rate video stream, with the frame rate hardcoded by API parameters, leaving the model no say in the matter.

Now the model decides for itself which segment to focus on, how quickly to process it, and whether to analyze the video, listen to the audio, or directly read the transcribed text.

Google

This isn't the first time Google has done this.

The previous agentic vision combined code execution with Gemini's native image understanding, allowing the model to write code autonomously to zoom, crop, and compare images.

Now it’s the video’s turn—same idea, just swapped the tool for the native video tool.

Four tasks that were once very difficult

The official blog also lists several high-difficulty use cases.

Sub-second clip retrieval. Previous models "viewed" videos by capturing only one frame per second. The problem is that many things pass by in less than a second.

We can now pinpoint these moments with precision under one second.

This means that “automatic editing” has become truly feasible for the first time: previously, editors had to manually scan the timeline frame by frame to determine where to cut, but now the model can first scan through and highlight potential cut points, allowing humans to make the final selections.

Google

Finding a needle in a haystack of long videos. Ask a complex question across hours of footage without burning through millions of tokens.

Anomaly detection. After the model identifies a specific time window, it can increase the frame rate and resample just that segment to capture fast motion and subtle visual defects. This is especially valuable for production line quality inspection and security surveillance, since anomalies typically occur only within those few seconds.

Count. Models used to consistently get questions like these wrong—how many times was an action repeated, how many different objects are in the scene—because the few missed frames just happened to contain the missing elements.

Google

The agent can finally view the video.

Video has always been the most expensive of all modalities.

Text is cheap, images are fine, but videos are large and long—one minute can consume as many tokens as a whole book.

So over the past two years, when everyone talked about agents, they were mostly discussing their ability to read, write, and use tools.

The issue is that an agent that primarily understands the world through text is only seeing a world that has been "recounted" by others.

It doesn’t know anything about what’s captured by cameras, recorded in meetings, or passing along production lines—things no one wants to spend time transcribing.

Now, this cost curve has been significantly reduced.

If an agent can independently process hours of surveillance footage, dozens of hours of courses, and hundreds of assets without exhausting the budget, its way of engaging with the world changes.

It no longer has to wait for someone to describe the visuals in words, nor does it have to choose between “seeing clearly” and “seeing accurately.”

Google was the first to break through.

AI Battle, Next Round

Over the past few days, the release schedules of the four parties almost coincided.

Claude 5.1 has just arrived, OpenAI Astra is already at the door, and Google, after months of anticipation, finally unveiled Gemini 3.8 Flash to compete.

After this update, Claude now focuses primarily on two areas: programming and scientific research.

The next generation of Astra will become a "continuous agent," and, in Otomo's words, its computer usage skills have reached human level.

Google

Gemini 3.8 Flash will also make a significant leap in programming.

Currently, OpenAI, Anthropic, Google, and SpaceX AI are all operating at full capacity.

Google

Musk previews Grok 4.7 to be released in ten days.

This chaos has just entered its most intense phase.

Reference materials:

https://x.com/Google/status/2094840983913402704

https://deepmind.google/blog/introducing-agentic-video-in-gemini/

This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.