Google Upgrades Gemini 3.8 Flash, Price Frozen Until 2026

icon币界网
Share
AI summary iconSummary
Google released Gemini 3.8 Flash and the 3.8 Flash Cyber variant for network upgrades, just three weeks after 3.7. The standard 3.8 Flash is now available via the Gemini API, Google AI Studio, Android Studio, and in select consumer products, while the Cyber version is accessible only to government agencies and infrastructure operators. Pricing remains frozen at $0.75 per million input tokens and $3.75 per million output tokens until the end of 2026, after which it will double. Complex tasks may require more tokens, increasing costs. Crypto price news indicates stable pricing for now, but changes are expected next year.
CoinDesk reports:

On September 2, Google released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber for cyber defense. This comes just three weeks after the launch of 3.7 Flash, marking the third Flash series update within six weeks. Google positions 3.8 as the most powerful Flash model for reasoning and coding to date; the standard version is now available via the Gemini API, Google AI Studio, Android Studio, Gemini Enterprise, and select consumer products. The Cyber version is exclusively available through the new Fairwind Program to trusted government agencies, critical infrastructure operators, and software maintainers. Both versions share the same foundational intelligence but differ in deployment permissions and security protections.

The introductory price for GPT-3.8 Flash remains $0.75 per million input tokens and $3.75 per million output tokens, unchanged from 3.7. However, this pricing applies only until December 31, 2026; starting January 1, 2027, the official rates will increase to $1.50 per million input tokens and $7.50 per million output tokens. Easier to overlook is that the model may perform additional reasoning steps and repeatedly invoke tools on complex tasks, and Google explicitly warns that high-effort tasks may consume more tokens. A stable per-token price does not guarantee that the total cost of a task will remain unchanged.

Flash begins competing for long tasks, with completion rate prioritized over speed.

The evidence provided by Google covers long-term coding, professional analysis, and multi-step reasoning. Flash outperforms most larger state-of-the-art models in the DeepSWE v1.1 long-term software engineering benchmark, achieving 54.9% on HLE-Verified. The company also cites financial and legal agent benchmarks to illustrate its effort to evolve Flash from a low-latency Q&A model into a work model capable of sustained planning, tool invocation, and delivering complete outcomes. Demonstrations such as hardware teardown visualizations, single-prompt generation of DOS-era maps, and 3D games emphasize that the model does not merely output code snippets, but can continuously inspect and refine its outputs in a loop.

However, the baseline score cannot be directly converted into enterprise productivity. Long-task success rates are often influenced by codebase size, dependency environments, test quality, tool permissions, and retry budgets. When the 3.8 “more diligent” strategy improves completion rates, it also increases execution time and token usage. Development teams must track per-call cost, total tokens per task, number of tool calls, number of retries before success, and manual rework time to determine whether an upgrade is truly more cost-effective. For latency- or budget-sensitive workloads, Google still recommends lowering the effort level or continuing to use 3.7 Flash; older versions are not immediately discontinued upon the release of new ones.

The standard version is now available to developers and enterprises, and Google AI Pro and Ultra users can also access it through the Gemini app, Search AI Mode, and Gemini in Sheets. Features, quotas, and regional availability may vary across different access points; therefore, “model released” should not imply that all users or all products have identical capabilities. In particular, workflows with autonomous tools depend on whether the application provides the model with the corresponding tools, permissions, and a recoverable execution environment.

The Cyber version is stronger and narrower, with 47.2% not being an auto-repair commitment.

Gemini 3.8 Flash Cyber focuses on vulnerability detection and patch generation. Google reports that it achieved a success rate of over 70% in internal vulnerability detection benchmarks covering 20 programming languages; in external CWE-Bench patch evaluation, it achieved a pass@1 score of 47.2%, nearing the 47.8% of a leading state-of-the-art model, but at lower cost. Internal usage by the Chrome security team showed that 3.8 Flash Cyber generated 2.6 times more correct vulnerability patches than the best large commercial model. Wiz also reported a 7.5 to 9.7 percentage point increase in recall on its internal penetration testing benchmark, with costs 2.3 to 5.2 times lower.

These figures are indicative but have limitations; the samples, prompts, and human evaluation methods for internal benchmarks were not fully disclosed in the announcement. A pass@1 of 47.2% also means that a single generation does not guarantee that more than half of the issues are correctly fixed. Google prioritizes “fixing” over “exploitation” and has added protections against chemical, biological, radiological, nuclear, and cyber abuse to the standard 3.8 model. The Cyber version, designed to meet professional defense requirements, employs more relaxed cybersecurity restrictions and is therefore not fully public, available only to trusted defenders within the Fairwind program.

What this release truly changes about the Flash series is its trade-off: instead of relying solely on low cost and speed to attract bulk requests, it now pursues higher completion rates for complex tasks through longer inference cycles. The most common mistake businesses make when directly replacing 3.7 with 3.8 is comparing only the price per million tokens. A more prudent approach is to test real tasks in tiers: limit effort for simple requests, allow more cycles for long tasks, and maintain human approval and isolated environments for security-sensitive tasks. Flash 3.8 is now available, and Cyber has entered a limited rollout; whether it can accomplish more production tasks at a lower total cost must be answered by each team’s own end-to-end data.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.