Gemini 3.8 Flash Review: Benchmark Leader but Real-World Letdown?
The author evaluates Google's Gemini 3.8 Flash model through hands-on coding, creative, and reasoning tests, finding it excels on benchmarks and cost-efficiency but falls short in agent capabilities, context retention, and practical tasks compared to rivals like Doubao and DeepSeek, while criticizing Google's product management and subscription value.
Background: Regret Over Gemini Annual Subscription
The author, an AI evaluator known as "Lao Zhang," recounts purchasing a 199-yuan annual Gemini Pro membership via Google's Antigravity IDE promotion, which originally offered access to Claude Opus 4.6. He describes subsequent issues: account suspension after using a third-party app to route Antigravity/Gemini models as APIs; drastic quota reductions for Claude Opus 4.6 (two questions consumed a week's allowance); persistent bugs in Antigravity IDE (memory leaks) and the deprecated Gemini desktop client; forced migration from Gemini CLI to "agy" with eligibility errors despite web/mobile access working; and outdated model lists still showing version 3.6.
Gemini 3.8 Flash on Paper: Strong Benchmarks, Weak Agent Skills
Official claims position Gemini 3.8 Flash as a cost-effective model with programming and vertical-domain performance nearing flagship levels at a fraction of the price. A programming cost-effectiveness chart (shown below) places it ahead of competitors. A specialized "Cyber" variant leads the CyberGym vulnerability-discovery benchmark.
However, community feedback suspects benchmark overfitting ("刷分"), as illustrated by a widely shared screenshot comparing benchmark scores vs. real-world performance.
Hands-On Testing: Mixed Results Across Tasks
1. Pelican on a Bicycle (Simon Willison's Classic Test)
The web version produced an acceptable static image for the "pelican riding a bicycle" prompt.
When asked to animate it as an SVG, the result was inferior to a previous attempt with Doubao-Seed-Evolving, which "秒杀" (instantly outperformed) Gemini.
2. Context Understanding Limitations
The web version struggled with context retention, as shown in a multi-turn conversation screenshot.
3. "Back View" Reading Comprehension + SVG Generation
Performance was comparable to Bisu 5.6's Sol model, each with strengths.
4. Rubik's Cube Recovery Simulation
The model handled the Rubik's cube recovery test smoothly.
5. PPT Creation
Results were poor: few images, overall quality far below DeepSeek-V4-Flash-vision.
6. High-Difficulty Test from X (Twitter) User
An external high-difficulty test showed Gemini 3.8 Flash underperforming not only Bisu and Opus but also K3.
Conclusion: Google's Resources vs. Product Execution
The author concludes that despite Google's vast talent, capital, and compute, the Gemini product experience is disappointing and "不争气" (fails to live up to potential).
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Old Zhang's AI Learning
AI practitioner specializing in large-model evaluation and on-premise deployment, agents, AI programming, Vibe Coding, general AI, and broader tech trends, with daily original technical articles.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
