Google Unveils Three Gemini Flash Models: Lower Token Use, Cheaper Batch Costs, and a Secure Pilot

Google released three Gemini Flash variants—3.6 Flash, 3.5 Flash‑Lite, and 3.5 Flash Cyber—each targeting different workloads, with the main model cutting token usage and inference steps, the Lite version reducing batch processing cost, and the Cyber version offering a controlled, security‑focused pilot.

ShiZhen AI
ShiZhen AI
ShiZhen AI
Google Unveils Three Gemini Flash Models: Lower Token Use, Cheaper Batch Costs, and a Secure Pilot

Google launches three Gemini Flash models

The announcement introduces Gemini 3.6 Flash as the new flagship model for coding, knowledge work, multimodal, and Agent tasks, optimized for shorter output, fewer reasoning steps, and fewer tool calls. Gemini 3.5 Flash‑Lite focuses on high‑frequency document, search, translation, and classification tasks, while Gemini 3.5 Flash Cyber is a security‑oriented variant deployed through CodeMender for vulnerability discovery and remediation.

Model positioning and pricing

Gemini 3.6 Flash – input $1.50 per M tokens, output $7.50 per M tokens.

Gemini 3.5 Flash‑Lite – input $0.30 per M tokens, output $2.50 per M tokens.

Gemini 3.5 Flash Cyber – pricing not publicly disclosed; limited to government and trusted partners via CodeMender.

All three models accept text, image, audio, and video inputs, support up to 1 M token context windows, and can generate up to 64 K tokens. Model identifiers in the API are gemini-3.6-flash and gemini-3.5-flash-lite.

3.6 Flash cuts redundant steps

Google cites data from Artificial Analysis showing that 3.6 Flash reduces output tokens by an average of 17% compared with 3.5 Flash, and up to 65% on the DeepSWE benchmark. It also trims unnecessary reasoning steps, tool calls, and code‑modification loops, which directly lowers latency and billing for long‑running Agent workloads.

Official benchmark figures show improvements such as DeepSWE (37% → 49%), MLE‑Bench (49.7% → 63.9%), OSWorld‑Verified (78.4% → 83.0%), and GDPval‑AA v2 (1349 → 1421). These gains are specific to coding, machine‑learning, knowledge‑work, and computer‑operation tasks and do not imply universal superiority across all inference tasks.

Flash‑Lite prioritizes throughput and cost

Flash‑Lite is designed for massive batch workloads where speed and per‑token price outweigh single‑query brilliance. Google reports an output rate of roughly 350 tokens per second, making it suitable for document processing, product information extraction, search, translation, and classification.

In official benchmarks, Flash‑Lite lifts Terminal‑Bench 2.1 from 31% to 54%, GDM‑MRCR v2 from 60.1% to 72.2%, and OSWorld‑Verified to 74%.

Flash Cyber focuses on secure code handling

Built on the 3.5 Flash base, Cyber undergoes security‑focused training and is orchestrated by CodeMender agents to discover, verify, and fix software vulnerabilities. In the CyberGym and Big Sleep evaluations, Cyber scores rise from 36%/42% (baseline models) to 72%.

Access is limited to government and trusted partners because an automated vulnerability‑finding model carries dual‑use risks.

Early community feedback

Within hours of release, users reported token reductions of 37% and latency drops of 60% for 3.6 Flash. Another tester saw execution time fall from 73 s to 34 s while token count dropped from 112,421 to 68,483, maintaining perfect task accuracy.

Conversely, some researchers observed a slight accuracy dip on an induction benchmark (7.8% vs. 10.9% for 3.5 Flash) and noted that performance gains may be benchmark‑specific noise.

Practical migration guidance

For teams using coding agents, computer‑operation, or multimodal document analysis, conduct an A/B test of 3.6 Flash, tracking success rate, total tokens, tool‑call count, and end‑to‑end latency. Prefer Flash‑Lite for high‑volume extraction, translation, or classification tasks. The Cyber variant is currently irrelevant for most developers and not slated for public API release.

When migrating to 3.6 Flash, focus on four metrics: success rate, total token consumption, number of tool calls, and overall latency, as early community results show the model is not a universal upgrade.

Google simultaneously releases three Gemini Flash models
Google simultaneously releases three Gemini Flash models
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AgentGoogleBenchmarkGeminitoken costAI modelsFlash
ShiZhen AI
Written by

ShiZhen AI

Tech blogger with over 10 years of experience at leading tech firms, AI efficiency and delivery expert focusing on AI productivity. Covers tech gadgets, AI-driven efficiency, and leisure— AI leisure community. 🛰 szzdzhp001

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.