Google's Triple Gemini Launch: 3.6 Flash, Flash-Lite, Cyber & Gemini 4 Pre-training Begins
Google DeepMind releases three specialized Gemini models — 3.6 Flash for token-efficient reasoning, 3.5 Flash-Lite for high-speed low-cost volume tasks, and 3.5 Flash Cyber for vulnerability detection — while confirming aggressive pre-training for Gemini 4, signaling a push to make production AI agents faster, cheaper, and more capable.
Google DeepMind announced three new Gemini models in a single release, each targeting distinct production workloads, while simultaneously confirming that pre-training for the next-generation Gemini 4 has begun.
Gemini 3.6 Flash: Token-Efficient Main Model
Positioned as the flagship of the trio, Gemini 3.6 Flash emphasizes reduced token consumption without sacrificing capability. According to Artificial Analysis Index benchmarks, it uses 17% fewer output tokens than its predecessor, 3.5 Flash. On the DeepSWE coding benchmark, token savings reach up to 65%. The model also requires fewer reasoning steps and tool calls to complete multi-step tasks, translating to lower latency and cost.
Pricing is set at $1.50 per million input tokens and $7.50 per million output tokens, undercutting 3.5 Flash. Despite the lower cost, 3.6 Flash improves on key benchmarks:
DeepSWE (coding): 37% → 49%
MLE Bench (ML research): 49.7% → 63.9%
OSWorld-Verified (computer control): 78.4% → 83%
GDPVal-AA v2 (knowledge work): 1,421 tasks completed, ~70 more than previous generation
Demonstrations show 3.6 Flash orchestrating multi-agent workflows for complex code migration, generating 3D texture-extraction tools via Gemini Canvas, and acting as an AI interaction designer for immersive theme studios.
Gemini 3.5 Flash-Lite: Speed and Volume Optimized
Flash-Lite targets high-throughput, cost-sensitive scenarios such as bulk document processing and agent-driven search. It delivers 350 output tokens per second — the fastest in the 3.5 series — at $0.30 per million input tokens and $2.50 per million output tokens.
Despite its lightweight profile, Flash-Lite outperforms the larger Gemini 3 Flash on several benchmarks:
SWE-Bench Pro: 54.2% vs. 49.6%
OSWorld-Verified: 74.0% vs. 65.1%
The suggested architecture pairs 3.6 Flash as the "planner" decomposing tasks with multiple Flash-Lite instances as parallel "workers." A demo generated 25 production-ready web designs in rapid succession. Flash-Lite is already live in the Gemini App and slated for integration into Google Search.
Gemini 3.5 Flash Cyber: Specialized Vulnerability Agent
Flash Cyber is a dedicated cybersecurity model focused on vulnerability discovery and remediation. Integrated into the CodeMender agent, it coordinates multiple Cyber-specialized agents to achieve state-of-the-art results on the CyberGym benchmark at lower cost than larger models. The model addresses the growing gap where AI-driven vulnerability detection outpaces traditional patching pipelines. Access is currently restricted and not available to the general public.
Gemini 4 Pre-training Underway
Google confirmed it has started what it describes as its "most aggressive pre-training run ever," targeting Gemini 4. Based on the company's typical six-month training cadence, a release could arrive by year-end. The article notes that while the three new models deliver immediate cost and efficiency gains — lower token usage, lower pricing, lower per-agent run cost — both the still-unreleased 3.5 Pro flagship and Gemini 4 remain "futures" for now.
Reference: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/?utm_source=tw&utm_medium=social&utm_campaign=og
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Top Architect
Top Architect focuses on sharing practical architecture knowledge, covering enterprise, system, website, large‑scale distributed, and high‑availability architectures, plus architecture adjustments using internet technologies. We welcome idea‑driven, sharing‑oriented architects to exchange and learn together.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
