Google DeepMind Unveils Three New Gemini Models: Lower Token Use, Higher Quality, Same Cost

Google DeepMind released three Gemini models—3.6 Flash, 3.5 Flash‑Lite and 3.5 Flash Cyber—offering up to 65% token savings, nearly double generation speed, improved benchmark scores and enhanced security while keeping pricing stable, with the flagship 3.5 Pro still delayed.

AI Engineering
AI Engineering
AI Engineering
Google DeepMind Unveils Three New Gemini Models: Lower Token Use, Higher Quality, Same Cost

Google announced three new Gemini models—3.6 Flash, 3.5 Flash‑Lite and 3.5 Flash Cyber—highlighting 3.6 Flash as a faster, cheaper, and more token‑efficient successor to the previous generation.

Pricing for 3.6 Flash drops the output cost from $9 to $7.50 per million tokens, while the input price remains $1.50. More significant is the token reduction: Google reports a 17% average decrease in output tokens on the Artificial Analysis Index, with up to 65% savings on certain DeepSWE coding tests, translating to a 30% lower bill ($727 vs $1,041 for the same workload).

Speed also improves markedly; 3.6 Flash generates 304 tokens/s compared with 165 tokens/s for 3.5 Flash, almost a two‑fold increase that benefits multi‑step reasoning and tool‑calling agent workflows.

Benchmark results show clear gains:

DeepSWE: 49% vs 37% (more precise code edits, fewer loops)

MLE‑Bench: 63.9% vs 49.7% (substantial uplift on ML research tasks)

OSWorld‑Verified: 83.0% vs 78.4% (enhanced Computer Use capability, now built into Gemini API and Enterprise)

GDPval‑AA v2: 1,421 vs 1,349 (stronger knowledge‑work performance)

The knowledge cutoff has been extended from January 2025 to March 2026, allowing the model to reference events from the past year.

Security enhancements include stronger defenses against CBRN threats and network‑attack abuse, improved resistance to jailbreak attempts, and a lower refusal rate for benign queries.

In head‑to‑head comparisons, 3.6 Flash leads on OSWorld and long‑context tests, while GPT‑5.6 Luna outperforms on DeepSWE and Terminal‑bench, Grok 4.5 leads on SWE‑Bench Pro, and Claude Sonnet 5 scores highest on MLE‑Bench and GDPval‑AA v2. Overall, most models are close, and selection depends on task and price.

3.6 Flash is already available in GitHub Copilot and has been integrated into the Gemini app, Google AI Studio, Android Studio, and the Antigravity agent‑development platform launched at I/O.

3.5 Flash‑Lite targets ultra‑low‑cost scenarios with a $0.30/$2.50 per‑million‑token pricing and a generation speed of 350 tokens/s, the fastest in the 3.5 series. It surpasses the prior 3.1 Flash‑Lite and even the original 3 Flash on several benchmarks:

SWE‑Bench Pro: 54.2% vs 49.6% (3 Flash)

OSWorld‑Verified: 74.0% vs 65.1% (3 Flash)

Terminal‑Bench 2.1: 54% vs 31% (3.1 Flash‑Lite)

GDM‑MRCR v2 (long‑context): 72.2% vs 60.1%

GDPval‑AA v2: 1,140 vs 642

Developers can choose different reasoning levels—minimal, low, or high—based on latency and cost requirements, with Computer Use offered as a built‑in tool. Early adopters such as Ashler, Palo Alto Networks and Ramp are already using the model for high‑throughput data processing, and Flash‑Lite will soon appear in Google Search.

3.5 Flash Cyber is a security‑focused variant fine‑tuned for vulnerability discovery and remediation. In the CyberGym benchmark it reaches 83.2%, comparable to GPT‑5.6 Sol (83.6%) and Mythos 5 (83.8%). Access is limited to governments and trusted partners via the CodeMender security agent.

The anticipated Gemini 3.5 Pro remains unavailable. Bloomberg reports internal testing fell short of expectations, and after a retraining effort in June the model still did not meet targets. It is currently being tested with a few partners and U.S. government entities, with no public release date announced.

Google confirmed that work on Gemini 4 has begun, representing the largest pre‑training effort to date.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LLMSecurityBenchmarkAI modelGeminiDeepMindToken efficiency
AI Engineering
Written by

AI Engineering

Focused on cutting‑edge product and technology information and practical experience sharing in the AI field (large models, MLOps/LLMOps, AI application development, AI infrastructure).

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.