First‑hand Benchmark: DeepSeek V4 Flash Across Claude Code, Codex, OpenCode, and Oh My Pi

A Composio test runs DeepSeek V4 Flash through four agent harnesses on 30 real multi‑step tasks integrating tools like Gmail and GitHub, revealing differences of up to three successful tasks, three‑fold cost variation, and 2.2× speed changes, with no single harness dominating all metrics.

PaperAgent
PaperAgent
PaperAgent
First‑hand Benchmark: DeepSeek V4 Flash Across Claude Code, Codex, OpenCode, and Oh My Pi

DeepSeek V4 Flash‑0731 was released and open‑sourced, featuring 284 B parameters with only 13 B active, and its agent benchmark surpasses V4‑Pro‑Preview.

Composio performed a test: they fed DeepSeek V4 Flash into four agent harnesses—Claude Code, Codex, OpenCode, and Oh My Pi—and ran the same 30 agentic tasks. The success rate differed by three tasks, cost differed by roughly three‑fold, and speed differed by 2.2×.

The 30 tasks are real multi‑step workflows that integrate online tools such as Gmail, Sheets, Airtable, GitHub, Slack, Calendar, Notion, and PagerDuty, covering accounting, auditing, batch editing, and synchronization. A task is considered passed only when all its fixed checks succeed.

Diagram
Diagram

Three dimensions, three different winners

Success rate (out of 30):

Oh My Pi: 17/30 (highest)

Claude Code: 16/30

Codex: 16/30

OpenCode: 14/30 (lowest)

Success rate comparison
Success rate comparison

Most counter‑intuitive: for seven tasks the pass/fail outcome depends entirely on which harness is used—same model, different wrapper changes the result.

Cost per successful task (estimated from API pricing):

OpenCode: $0.073 (cheapest)

Codex: $0.081

Oh My Pi: $0.103

Claude Code: $0.195 (most expensive, about 2.7 × the cheapest)

Cost comparison
Cost comparison

All four harnesses stay below $0.20 per successful task, but the author warns that DeepSeek’s price increase could break this cost pattern.

Median end‑to‑end time per task:

Claude Code: 122.7 s (fastest)

OpenCode: 129.7 s

Codex: 245.0 s

Oh My Pi: 272.4 s (slowest)

Speed comparison
Speed comparison

No all‑round winner

Comparing the three dimensions shows that the fastest is also the most expensive, and the highest success rate is the slowest. Claude Code is fastest but most costly; Oh My Pi has the highest success rate but is slowest; OpenCode is cheapest but has the lowest success rate. No harness dominates all three metrics.

Conclusion: the choice of harness changes the outcome of three tasks, nearly triples cost, and alters speed by 2.2×. Therefore, select the harness based on the metric you value most rather than a single overall ranking.

Composio - We ran DeepSeek V4 Flash through 4 agent harnesses (Claude Code, Codex, OpenCode, Oh My Pi) on 30 agentic tasks
 https://x.com/composio/status/2085330847951970801
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

PerformancebenchmarkCost AnalysisCodexClaude CodeAgent HarnessDeepSeek V4 FlashOh My Pi
PaperAgent
Written by

PaperAgent

Daily updates, analyzing cutting-edge AI research papers

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.