Gemini 3.8 Flash Review: Benchmark Leader but Real-World Letdown?

The author evaluates Google's Gemini 3.8 Flash model through hands-on coding, creative, and reasoning tests, finding it excels on benchmarks and cost-efficiency but falls short in agent capabilities, context retention, and practical tasks compared to rivals like Doubao and DeepSeek, while criticizing Google's product management and subscription value.

Old Zhang's AI Learning
Old Zhang's AI Learning
Old Zhang's AI Learning
Gemini 3.8 Flash Review: Benchmark Leader but Real-World Letdown?

Background: Regret Over Gemini Annual Subscription

The author, an AI evaluator known as "Lao Zhang," recounts purchasing a 199-yuan annual Gemini Pro membership via Google's Antigravity IDE promotion, which originally offered access to Claude Opus 4.6. He describes subsequent issues: account suspension after using a third-party app to route Antigravity/Gemini models as APIs; drastic quota reductions for Claude Opus 4.6 (two questions consumed a week's allowance); persistent bugs in Antigravity IDE (memory leaks) and the deprecated Gemini desktop client; forced migration from Gemini CLI to "agy" with eligibility errors despite web/mobile access working; and outdated model lists still showing version 3.6.

Gemini 3.8 Flash on Paper: Strong Benchmarks, Weak Agent Skills

Official claims position Gemini 3.8 Flash as a cost-effective model with programming and vertical-domain performance nearing flagship levels at a fraction of the price. A programming cost-effectiveness chart (shown below) places it ahead of competitors. A specialized "Cyber" variant leads the CyberGym vulnerability-discovery benchmark.

Programming cost-effectiveness chart for Gemini 3.8 Flash
Programming cost-effectiveness chart for Gemini 3.8 Flash
CyberGym benchmark results for Gemini 3.8 Flash Cyber version
CyberGym benchmark results for Gemini 3.8 Flash Cyber version

However, community feedback suspects benchmark overfitting ("刷分"), as illustrated by a widely shared screenshot comparing benchmark scores vs. real-world performance.

Community screenshot suggesting benchmark gaming
Community screenshot suggesting benchmark gaming

Hands-On Testing: Mixed Results Across Tasks

1. Pelican on a Bicycle (Simon Willison's Classic Test)

The web version produced an acceptable static image for the "pelican riding a bicycle" prompt.

Pelican on bicycle static output
Pelican on bicycle static output

When asked to animate it as an SVG, the result was inferior to a previous attempt with Doubao-Seed-Evolving, which "秒杀" (instantly outperformed) Gemini.

2. Context Understanding Limitations

The web version struggled with context retention, as shown in a multi-turn conversation screenshot.

Context understanding failure example
Context understanding failure example

3. "Back View" Reading Comprehension + SVG Generation

Performance was comparable to Bisu 5.6's Sol model, each with strengths.

Back View reading comprehension and SVG output
Back View reading comprehension and SVG output
Comparison with Bisu 5.6 Sol
Comparison with Bisu 5.6 Sol

4. Rubik's Cube Recovery Simulation

The model handled the Rubik's cube recovery test smoothly.

5. PPT Creation

Results were poor: few images, overall quality far below DeepSeek-V4-Flash-vision.

PPT generation output comparison
PPT generation output comparison

6. High-Difficulty Test from X (Twitter) User

An external high-difficulty test showed Gemini 3.8 Flash underperforming not only Bisu and Opus but also K3.

Conclusion: Google's Resources vs. Product Execution

The author concludes that despite Google's vast talent, capital, and compute, the Gemini product experience is disappointing and "不争气" (fails to live up to potential).

Final commentary image
Final commentary image
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

model comparisoncoding assistantsAI model benchmarkingcybersecurity AIGemini 3.8 FlashGoogle AI productshands-on testingsubscription regret
Old Zhang's AI Learning
Written by

Old Zhang's AI Learning

AI practitioner specializing in large-model evaluation and on-premise deployment, agents, AI programming, Vibe Coding, general AI, and broader tech trends, with daily original technical articles.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.