DeepSeek V4 Pro: First‑hand Test After Rumored Withdrawal – How It Stacks Up

The author investigates DeepSeek V4 Pro after a rumored release and sudden withdrawal, examining API errors, benchmark scores on AutomationBench and CyberGym, pricing versus V4 Flash, scaling limits, three practical demo cases, and the broader open‑source LLM landscape in China.

Baobao Algorithm Notes
Baobao Algorithm Notes
Baobao Algorithm Notes
DeepSeek V4 Pro: First‑hand Test After Rumored Withdrawal – How It Stacks Up

On the night of August 12 a link appeared in DeepSeek’s official group, followed by rumors that the announcement was retracted later that afternoon; the website’s news post disappeared, yet the API and pricing docs remained until the API began returning errors at 18:00. The author decided to let the "bullet fly" and test the rumored "official" V4 Pro release.

Benchmark images show that V4 Pro’s agent performance approaches Claude’s Fable 5 and engages Opus 4.8 in a back‑and‑forth dialogue. On two specific tests—AutomationBench (measuring an AI agent’s ability to perform real work) and CyberGym (evaluating AI in realistic network‑security offense‑defense scenarios)—V4 Pro surpasses Fable 5.

Pricing has not changed: V4 Pro costs $6 per million output tokens and $3 for input cache misses, three times the $2/$1 rates of V4 Flash. Despite the three‑fold price increase, the author argues the performance gain is far less than three‑fold; for everyday tasks such as coding, writing, spreadsheet work, or simple web pages, V4 Flash already performs adequately.

AA Index data released the next morning confirms that V4 Pro’s improvement over V4 Flash is modest. The author interprets this as a sign that scaling the model’s parameters (1.6 T for Pro vs. 284 B for Flash, a >5× increase) yields diminishing returns when post‑training data quality is similar.

Three practical case studies illustrate the model’s capabilities:

A web‑based Linux terminal that supports 25 commands, pipelines (e.g., ls | grep txt), redirection ( echo hello > file.txt), and even a minimal Vim editor.

A 3D dinosaur game built with Three.js, featuring desert navigation, obstacle avoidance, day/night cycles, and a switch from third‑person to side‑view perspective upon request.

A web‑based spreadsheet with a formula engine implementing over 30 functions (SUM, AVERAGE, VLOOKUP, etc.), supporting relative and range references, copy‑paste, multi‑sheet handling, CSV export, and persistent local storage.

The author reflects on three takeaways: cost‑effectiveness, scaling limits, and the broader market dynamics. While V4 Pro’s larger parameter count brings some improvement, the marginal gains suggest that both parameter scaling and test‑time compute are approaching convergence. The Chinese open‑source LLM scene—DeepSeek, GLM, and Kimi—has entered a rapid “rotation” where each model briefly leads, all while being fully open source, a development the author deems more significant than any single benchmark score.

In conclusion, V4 Pro is strong enough to stand alongside leading overseas models, but its true advantage lies in the contributions of the domestic open‑source community rather than raw size or price.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agentsLLMopen-sourceDeepSeekparameter scalingBenchmarkV4 Pro
Baobao Algorithm Notes
Written by

Baobao Algorithm Notes

Author of the BaiMian large model, offering technology and industry insights.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.