Google Unveils Gemini 3.6 Flash, 3.5 Flash‑Lite, and Flash‑Cyber – Gemini 4 Pre‑training Begins

Google DeepMind released three Gemini variants—3.6 Flash, 3.5 Flash‑Lite, and 3.5 Flash‑Cyber—each promising token savings, higher speed, or security focus, while also announcing an aggressive pre‑training run for the upcoming Gemini 4, signaling a push for cheaper, faster AI agents.

Top Architect
Top Architect
Top Architect
Google Unveils Gemini 3.6 Flash, 3.5 Flash‑Lite, and Flash‑Cyber – Gemini 4 Pre‑training Begins

Gemini 3.6 Flash – Token‑saving powerhouse

The flagship model emphasizes token efficiency, using 17% fewer output tokens than the previous 3.5 Flash and achieving up to 65% savings on the DeepSWE coding benchmark. Pricing is $1.5 per million input tokens and $7.5 per million output tokens, cheaper than its predecessor. Benchmark improvements include DeepSWE 37% → 49%, MLE Bench 49.7% → 63.9%, OSWorld‑Verified 78.4% → 83%, and GDPVal‑AA v2 scoring 1,421 points, about 70 points higher than the prior model. Demo videos show multi‑agent orchestration, 3D workflow integration via Gemini Canvas, and AI‑driven interaction design.

Gemini 3.5 Flash‑Lite – Speed‑focused lightweight

This variant targets raw throughput, reaching 350 tokens per second—the fastest in the 3.5 series—and costs only $0.3 per million input tokens and $2.5 per million output tokens. It is positioned for high‑frequency document processing and agent‑search scenarios. In head‑to‑head tests it outperforms the larger Gemini 3 Flash: SWE‑Bench Pro 54.2% vs 49.6% and OSWorld‑Verified 74.0% vs 65.1%. Flash‑Lite is already available in the Gemini App and will be rolled out to Google Search.

Gemini 3.5 Flash‑Cyber – Security‑oriented model

Designed exclusively for vulnerability detection and remediation, Flash‑Cyber integrates into the CodeMender security agent. It leverages multiple Cyber agents to achieve state‑of‑the‑art results on the CyberGym benchmark while incurring lower cost than larger models. The model is not yet publicly accessible.

Gemini 4 – The next generation

Google confirmed that an “aggressive” pre‑training run for Gemini 4 has started, aiming for a release by year‑end following the company’s typical six‑month training cadence. While Gemini 4 remains a future flagship, the three simultaneous releases focus on reducing token consumption, lowering cost, and improving speed for production AI agents.

For users who pay per token, the immediate benefit is a clear reduction in both token usage and per‑run cost, whereas the promised Gemini 4 represents a longer‑term upgrade path.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI Agentslarge language modelbenchmarkGeminiGoogle AIsecurity modelToken efficiency
Top Architect
Written by

Top Architect

Top Architect focuses on sharing practical architecture knowledge, covering enterprise, system, website, large‑scale distributed, and high‑availability architectures, plus architecture adjustments using internet technologies. We welcome idea‑driven, sharing‑oriented architects to exchange and learn together.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.