Google Launches Three Gemini Models, Starts Gemini 4 Training
Google DeepMind released three new Gemini models—3.6 Flash with 65% token reduction, 3.5 Flash-Lite for high-speed low-cost processing, and 3.5 Flash Cyber for vulnerability detection—while simultaneously beginning aggressive pre-training for Gemini 4, signaling continued rapid advancement in AI agent capabilities and cost reduction.
Google DeepMind's Triple Gemini Release
Google DeepMind announced three new Gemini models in a single release, each targeting different production AI agent needs: a stronger main model, a faster lightweight variant, and a specialized security model. The company also confirmed it has begun its "most aggressive pre-training ever" for the next-generation Gemini 4.
Gemini 3.6 Flash: Main Model with 65% Token Reduction
Gemini 3.6 Flash is positioned as the new workhorse model. Its headline improvement is dramatic token efficiency: according to Artificial Analysis Index, it uses 17% fewer output tokens than 3.5 Flash overall, and up to 65% fewer on the DeepSWE coding benchmark. It also requires fewer reasoning steps and tool calls to complete multi-step tasks, making it both faster and cheaper.
Pricing dropped to $1.50 per million input tokens and $7.50 per million output tokens, cheaper than 3.5 Flash. Despite lower cost, benchmark scores improved significantly:
DeepSWE (coding): 37% → 49%
MLE Bench (ML research): 49.7% → 63.9%
OSWorld-Verified (computer operation): 78.4% → 83%
GDPVal-AA v2 (knowledge work): 1,421 tasks completed, 70+ more than predecessor
Demos showed 3.6 Flash orchestrating multi-agent code migration, generating 3D workflows via Gemini Canvas, and acting as an "AI interaction designer" for immersive theme studios.
Gemini 3.5 Flash-Lite: Speed and Cost for High-Volume Workloads
Flash-Lite targets high-frequency, high-volume scenarios like massive document processing and agent search. It outputs 350 tokens/second—the fastest in the 3.5 series—at $0.30 per million input tokens and $2.50 per million output tokens.
Surprisingly, this lightweight model outperformed the larger Gemini 3 Flash on several benchmarks:
SWE-Bench Pro: 54.2% vs 49.6%
OSWorld-Verified: 74.0% vs 65.1%
The suggested architecture pairs 3.6 Flash as the "brain" for task decomposition with multiple Flash-Lite instances as "workers" for parallel execution. A demo generated 25 production-ready web designs almost instantly. Flash-Lite is already live in the Gemini App and rolling into Google Search.
Gemini 3.5 Flash Cyber: Specialized Vulnerability Detection
Flash Cyber is a dedicated cybersecurity model integrated into the CodeMender code security agent. Multiple Cyber agents collaborate to find and fix vulnerabilities, achieving state-of-the-art on the CyberGym benchmark at lower cost than larger models. The model addresses a growing gap: AI now discovers vulnerabilities faster than existing systems can patch them. Access is currently restricted; not available to general users.
Gemini 4 Pre-Training Underway
On the same day, Google confirmed it has started its "most aggressive pre-training ever" targeting Gemini 4. Following Google's typical six-month training cycle, Gemini 4 could appear by year-end. The article notes that for developers paying for tokens and building agents today, the immediate value is cost reduction: lower prices, lower token consumption, and lower per-agent-run costs. The flagship models remain "futures" — practical, affordable models are what can be deployed now.
Reference: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/?utm_source=tw&utm_medium=social&utm_campaign=og
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Top Architect
Top Architect focuses on sharing practical architecture knowledge, covering enterprise, system, website, large‑scale distributed, and high‑availability architectures, plus architecture adjustments using internet technologies. We welcome idea‑driven, sharing‑oriented architects to exchange and learn together.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
