Google Unveils Gemini 3.6 Flash, 3.5 Flash‑Lite, and Flash‑Cyber – Gemini 4 Pre‑training Begins
Google DeepMind released three Gemini variants—3.6 Flash, 3.5 Flash‑Lite, and 3.5 Flash‑Cyber—each promising token savings, higher speed, or security focus, while also announcing an aggressive pre‑training run for the upcoming Gemini 4, signaling a push for cheaper, faster AI agents.
Gemini 3.6 Flash – Token‑saving powerhouse
The flagship model emphasizes token efficiency, using 17% fewer output tokens than the previous 3.5 Flash and achieving up to 65% savings on the DeepSWE coding benchmark. Pricing is $1.5 per million input tokens and $7.5 per million output tokens, cheaper than its predecessor. Benchmark improvements include DeepSWE 37% → 49%, MLE Bench 49.7% → 63.9%, OSWorld‑Verified 78.4% → 83%, and GDPVal‑AA v2 scoring 1,421 points, about 70 points higher than the prior model. Demo videos show multi‑agent orchestration, 3D workflow integration via Gemini Canvas, and AI‑driven interaction design.
Gemini 3.5 Flash‑Lite – Speed‑focused lightweight
This variant targets raw throughput, reaching 350 tokens per second—the fastest in the 3.5 series—and costs only $0.3 per million input tokens and $2.5 per million output tokens. It is positioned for high‑frequency document processing and agent‑search scenarios. In head‑to‑head tests it outperforms the larger Gemini 3 Flash: SWE‑Bench Pro 54.2% vs 49.6% and OSWorld‑Verified 74.0% vs 65.1%. Flash‑Lite is already available in the Gemini App and will be rolled out to Google Search.
Gemini 3.5 Flash‑Cyber – Security‑oriented model
Designed exclusively for vulnerability detection and remediation, Flash‑Cyber integrates into the CodeMender security agent. It leverages multiple Cyber agents to achieve state‑of‑the‑art results on the CyberGym benchmark while incurring lower cost than larger models. The model is not yet publicly accessible.
Gemini 4 – The next generation
Google confirmed that an “aggressive” pre‑training run for Gemini 4 has started, aiming for a release by year‑end following the company’s typical six‑month training cadence. While Gemini 4 remains a future flagship, the three simultaneous releases focus on reducing token consumption, lowering cost, and improving speed for production AI agents.
For users who pay per token, the immediate benefit is a clear reduction in both token usage and per‑run cost, whereas the promised Gemini 4 represents a longer‑term upgrade path.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Top Architect
Top Architect focuses on sharing practical architecture knowledge, covering enterprise, system, website, large‑scale distributed, and high‑availability architectures, plus architecture adjustments using internet technologies. We welcome idea‑driven, sharing‑oriented architects to exchange and learn together.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
