Google's Gemini 3.8 Flash targets agents, Cyber twin finds 13-year-old Chrome bug
Two new Flash models: one for agentic work, one for cybersecurity, with Flash Cyber already patching Chrome and finding a decade-old flaw.

Google CEO Sundar Pichai announced Gemini 3.8 Flash and Flash Cyber, the company's most capable cybersecurity model. The release signals a strategic push into agentic AI and defensive security, with Flash Cyber already securing Google's own code.
Google's Gemini 3.8 Flash isn't just another model drop - it's a two-pronged assault on the AI frontier. On Wednesday, CEO Sundar Pichai unveiled two variants: a standard Flash built as a "workhorse" for agentic tasks, software development, and multi-step reasoning, and Flash Cyber, which Google calls its "most capable" cybersecurity model. The latter already found a vulnerability that had lurked in Chromium and Chrome for 13 years - a "very subtle bug" that dozens, if not hundreds, of engineers had reviewed but never flagged, according to Doug Turner, engineering director for Chrome. That discovery, shared in a Google video, is the kind of proof point that makes enterprises sit up and take notice.
The 13-year-old bug is just one data point in a broader story. Flash Cyber scored 86.2% on the CyberGym benchmark and 47.2% on CWE-Bench, which tests AI patching abilities. In an internal Google benchmark, it achieved a more than 70% success rate discovering vulnerabilities across 20 programming languages. Pichai touted "significant leaps" over the previous 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. The model also outperformed many larger frontier models on the DeepSWE coding benchmark at far lower cost - a direct challenge to the assumption that bigger always means better.
This is Google's third Flash release in six weeks, a cadence that signals an aggressive push to own the agentic AI layer. The standard 3.8 Flash is available now in Gemini Enterprise, with developers able to test it via the Gemini API, Google AI Studio, Antigravity, Android Studio, or Stitch. Pricing holds at $0.75 per million input tokens and $3.75 per million output tokens - the same introductory rate as 3.7 Flash. That's a deliberate move to undercut rivals on cost while offering adjustable "effort levels" so users can trade quality, latency, and token spend. As Google senior product director Tulsee Doshi and Gemini security lead Raluca Ada Popa wrote, "3.8 Flash works harder," showing "greater diligence" on complex tasks, sometimes using more tokens to maximize performance.
The model's benchmark results are impressive but not uniform. It landed at No. 14 on Arena.ai's Agent Arena, above DeepSeek-V4-Pro and a significant jump from 3.7 Flash's No. 32 spot. It debuted at No. 7 in Text Arena, ahead of Claude Opus 5. It improved over 3.7 Flash in multi-turn requests, writing, literature, language, longer queries, hard prompts, coding, instruction following, and business operations. On specialized benchmarks, it outperformed rivals in finance (Vals Finance Agent V2) and law (Harvey's Legal Agent Benchmark), and scored 54.9% on Humanity's Last Exam (HLE)-Verified, reflecting strong multi-step reasoning across math, science, and humanities.
But the real strategic weight sits with Flash Cyber. It's initially rolling out only to "trusted defenders" through Google's Fairwind Program, which prioritizes government authorities, critical-infrastructure operators, and other partners seeking advanced cyber defense. That gated access reflects the model's "more permissive set of mitigations" for cybersecurity safeguards - a careful balance between empowering defenders and preventing misuse in offensive cyber or CBRN domains. Google says Flash Cyber has undergone "rigorous training" and represents a "significant leap in prompt injection robustness." The company is already using it to secure its own code: it produced 2.6 times more correct patches in Chrome vulnerabilities versus much larger commercial models.
The economics are striking. Wiz, which Google acquired earlier this year for a historic $32 billion, reported that Flash Cyber had 7.5% to 9.7% higher recall of real-world vulnerabilities on an internal penetration testing benchmark at 2.3 to 5.2 times lower cost than leading frontier models. Google's Cloud Vulnerability Research found a critical foundational vulnerability in less than 2 hours with Flash Cyber - work that typically takes months. That speed matters because defenders are overwhelmed. As Popa put it, "In cybersecurity, attackers need only find one significant flaw over millions of lines of code. Defenders have to remove every one of those flaws to be able to defend against attackers." Turner described a "vulnerability apocalypse" in recent months due to generative AI, with a hockey-stick increase in reported vulnerabilities.
For executives, the takeaway is clear: AI is no longer just a productivity tool - it's becoming a defensive necessity. Flash Cyber's ability to find and patch vulnerabilities at scale, at a fraction of the cost of traditional methods, could reshape how companies approach security budgets. The 13-year-old bug is a reminder that human review has limits, and AI can see what we miss. But the gated rollout also signals that this power is dangerous in the wrong hands. Google's decision to prioritize vulnerability fixing over offensive exploitation - as Doshi and Popa emphasized - sets a tone for the industry. As AI agents become "incredibly skilled" at finding and exploiting flaws, the race between offense and defense is accelerating. Companies that wait to adopt these tools may find themselves on the wrong side of that curve.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology
BASF sues Apple over Face ID, dragging iPhone and iPad into Texas court
The world's largest chemical company claims dozens of Apple devices infringe its face authentication patents - and it chose a venue known for fast, plaintiff-friendly patent trials.
Uber's UK robotaxi debut: 15 self-driving cars, safety drivers inside
The ride-hailing giant's first UK autonomous fleet is a cautious pilot; here's what it signals for the robotaxi race.
Muse Spark 1.3 goes live as Zuckerberg promises open-weights release 'soon'
Meta ships a leaner, less chatty flagship model to its API while teasing an open-weights version, sharpening its challenge to GPT-5.6 and Claude.



