TradingKey - Google (GOOGL) released two new models on September 2, Eastern Time: Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. This is the third Flash series model released by Google in six weeks, continuing the high-speed and low-cost features of the previous Gemini 3.7 Flash while further enhancing code generation, complex reasoning, agentic tasks, and cybersecurity capabilities.
The former is positioned for software engineering, AI agents, and multi-step reasoning tasks in specialized fields; the latter focuses on vulnerability discovery and automated remediation, available only to vetted, trusted defenders.
Although the two models target different use cases, both are based on the same underlying intelligence core and repeatedly evaluate and optimize model outputs through long-running agent loops.
Gemini 3.8 Flash: Enhances Long-Horizon Coding and Agent Tasks
The primary upgrade in Gemini 3.8 Flash centers on handling complex, continuous tasks that require multi-turn tool calling. Google stated that in select tests, the model's performance approaches that of higher-cost flagship models.
In the DeepSWE v1.1 long-horizon software engineering benchmark, Gemini 3.8 Flash scored 73.7%, higher than Gemini 3.7 Flash's 65.3%, and also exceeding GPT-5.6 Sol's 72.7% and GPT-5.6 Terra's 69.6%, falling just slightly behind Claude Opus 5's 74.0%.

Source: Google
In specialized domain tasks, Gemini 3.8 Flash also demonstrated strong competitiveness. Among these, HLE-Verified covers multidisciplinary questions across science, technology, humanities, and professional knowledge, where Gemini 3.8 Flash also outperformed Claude Sonnet 5's 31.0%, GPT-5.6 Terra's 51.1%, Claude Opus 5's 54.4%, and GPT-5.6 Sol's 54.5%.
On pricing, Gemini 3.8 Flash maintains the initial pricing of Gemini 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens.
By comparison, input and output prices for Claude Opus 5 are $5 and $25, respectively; for GPT-5.6 Sol, $4 and $20; and for GPT-5.6 Terra, $2 and $12.
This means Gemini 3.8 Flash approaches or even surpasses larger models in select coding and reasoning benchmarks, but at a significantly lower API call cost. For developers looking to deploy code agents, automated analytics, and enterprise workflows, its cost-efficiency could serve as its key selling point.
Gemini 3.8 Flash Cyber: Focuses on Vulnerability Discovery and Automated Remediation
Gemini 3.8 Flash Cyber is another key model released this time, primarily targeting cybersecurity defense scenarios. Google stated that the model will not be broadly available, but will instead be provided to vetted, trusted defense teams through the new Fairwind program.
In the CyberGym vulnerability discovery benchmark commonly used in the cybersecurity industry, Gemini 3.8 Flash Cyber scored 86.2%, higher than Gemini 3.5 Flash Cyber's 77.5%, and also surpassing GPT-5.6 Sol's 83.6%, Mythos 5's 83.8%, and GPT-5.5-Cyber's 85.6%.
In internal real-world scenario testing covering 20 programming languages, Gemini 3.8 Flash Cyber achieved a vulnerability discovery success rate of 71.0%, higher than Gemini 3.7 Flash's 58.9% and Gemini 3.5 Flash Cyber's 46.6%.
In terms of automated remediation, Gemini 3.8 Flash Cyber achieved a pass@1 of 47.2% in the CWE-Bench test run by Collinear, close to the 47.8% of leading top-tier models, though Google highlighted its lower cost.
Find out more