Gemini 3.8 Flash Hits 71% on DeepSWE — And Its Cyber Twin Is Government-Only Again


Google released Gemini 3.8 Flash on September 2, 2026 — just three weeks after Gemini 3.7 Flash — alongside a restricted security variant called Gemini 3.8 Flash Cyber, which ships exclusively to vetted partners through a new access program called the Fairwind Program.
The headline metric is DeepSWE v1.1: 71.0% — a 5.7-point improvement over 3.7 Flash's 65.3% on the same benchmark, which measures autonomous end-to-end software engineering on real repositories. On HLE-Verified, which tests multi-step STEM and professional reasoning, 3.8 Flash scored 54.9%. Google also reports gains over 3.7 Flash on the Vals Finance Agent V2 and Harvey's Legal Agent Benchmark, suggesting post-training improvements that generalise across domain-specific agentic tasks, not just coding.
The architecture is unchanged in the ways that matter for developers: 3.8 Flash retains the 1-million-token context window, 64,000 output token limit, and the same three-level thinking effort control (low, medium, high) introduced in 3.7 Flash. Google says the performance gains come from further refinement of post-training on agentic trajectories, with reinforcement learning focused on tool-calling fidelity in iterative workflows.
Pricing is held at the same introductory rates as 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens, locked in through December 31, 2026, at which point standard pricing of $1.50/$7.50 per million tokens takes effect. The model is available now across the Gemini API, AI Studio, and Vertex AI.
Gemini 3.8 Flash Cyber is the more pointed announcement. It follows the access model Google established with Gemini 3.5 Flash Cyber in July — a specialized vulnerability-hunting model that does not ship to the public — but with an upgraded delivery mechanism. Rather than the earlier CodeMender pilot, Google is now formalizing restricted-access distribution under the Fairwind Program: a vetting-gated initiative for government authorities, critical national infrastructure operators, and software maintainers. The program integrates 3.8 Flash Cyber with the CodeMender harness for autonomous patch generation and validation, and includes specific safeguards against offensive-domain misuse that are not present in the standard variant.
Google describes 3.8 Flash Cyber as its most capable cybersecurity model to date, outperforming both 3.7 Flash Cyber and some larger commercial models on CyberGym and internal penetration testing benchmarks, though neither benchmark result nor an external audit has been published. The pattern — frontier cyber capability restricted behind a vetting program — now mirrors what Anthropic does with Mythos and what Microsoft did with MAI-Cyber-1-Flash's Project Perception preview. Controlled distribution at launch rather than guardrails-after-the-fact is clearly becoming the default posture for high-capability security models across the industry.


