DeepSeek V4 Pro vs. Grok 4.6: How New LLMs Challenge Top Closed‑Source Models

The newly released DeepSeek V4 Pro and Elon Musk’s Grok 4.6 deliver performance and cost metrics that rival or surpass leading closed‑source LLMs, with DeepSeek achieving up to 29‑fold cheaper token output and top scores on Agent, CyberGym, AutomationBench, Terminal‑Bench, and professional legal benchmarks, while Grok 4.6 matches GPT‑5.6 on the AA Intelligence Index and leads in workplace knowledge tests.

SuanNi
SuanNi
SuanNi
DeepSeek V4 Pro vs. Grok 4.6: How New LLMs Challenge Top Closed‑Source Models

DeepSeek announced the official release of its V4 Pro model, and Elon Musk’s X AI unveiled Grok 4.6 on the same day, positioning both models as direct competitors to leading closed‑source large language models.

Cost analysis shows that DeepSeek charges only ¥6 (≈$0.85) per million output tokens, far cheaper than GPT‑5.6 Sol ($30), Claude Opus 5 ($25), and Fable 5 ($50). Grok 4.6 costs $6 per million tokens, matching DeepSeek’s price advantage over GPT‑5.6.

DeepSeek V4 Pro performance : The model’s Agent capability improves dramatically. In the CyberGym security‑agent test it scores 83.3, surpassing Fable 5 (83.1) and Opus 4.8 (78.3). AutomationBench yields 31.8, beating Fable 5 (29.1) and Opus 4.8 (27.2). On Terminal‑Bench 2.1 DeepSeek reaches 87.9, just 0.1 behind Fable 5 (88.0) and ahead of Opus 4.8 (85.0). In the Human‑Level‑Exam (HLE) it scores 60.0, overtaking Opus 4.8 (57.9). Software‑engineering evaluation (DeepSWE) jumps from 12.8 in the preview to 62.7 in the official release, a 4.9‑fold increase, exceeding Opus 4.8’s 58.0. Toolathon‑Verified records 74.1, slightly below Opus 4.8 (76.2) and Kimi‑K3 (76.5).

Grok 4.6 performance : On the AA Intelligence Index it attains 61 points, equal to GPT‑5.6 Sol and only one point shy of Fable 5 (62), a 5‑point rise from Grok 4.5’s 56. In the code‑editor benchmark CursorBench it scores 69.9% versus GPT‑5.6’s 67.2%, trailing Fable 5 by 0.6 points. FrontierCode places Grok 4.6 at 61.3%, ahead of GPT‑5.6 (60.6%) and close to Fable 5 (64.9%). Workplace‑knowledge tests are all topped by Grok 4.6. GDPval‑AA v2 yields 1753 Elo, surpassing Fable 5 (1741) and GPT‑5.6 (1728). AA‑Briefcase records 1577 Elo, edging out Fable 5 (1574). In the professional legal benchmark Harvey LAB, Grok 4.6 achieves 15.8% versus Fable 5’s 11.3% and GPT‑5.6’s 2.5%.

Databricks engineer Ivan Zhou reported that, using the Dbrx Mosaic AI OfficeQA Pro V2 suite together with Databricks’ Genie Harness, Grok 4.6 attained the best scores in document understanding and data‑reasoning tasks.

Future outlooks: Musk hinted at a forthcoming Grok 4.7 that will outperform all existing models, while internal sources claim DeepSeek’s upcoming Harness will become the “strongest agent in the universe.” The rapid release of these high‑performing models compresses the competitive landscape, putting established providers such as OpenAI and Anthropic under pressure.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

aiLLMDeepSeekBenchmarkGrokcostagent performance
SuanNi
Written by

SuanNi

A community for AI developers that aggregates large-model development services, models, and compute power.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.