First‑hand Benchmark: DeepSeek V4 Flash Across Claude Code, Codex, OpenCode, and Oh My Pi
A Composio test runs DeepSeek V4 Flash through four agent harnesses on 30 real multi‑step tasks integrating tools like Gmail and GitHub, revealing differences of up to three successful tasks, three‑fold cost variation, and 2.2× speed changes, with no single harness dominating all metrics.
DeepSeek V4 Flash‑0731 was released and open‑sourced, featuring 284 B parameters with only 13 B active, and its agent benchmark surpasses V4‑Pro‑Preview.
Composio performed a test: they fed DeepSeek V4 Flash into four agent harnesses—Claude Code, Codex, OpenCode, and Oh My Pi—and ran the same 30 agentic tasks. The success rate differed by three tasks, cost differed by roughly three‑fold, and speed differed by 2.2×.
The 30 tasks are real multi‑step workflows that integrate online tools such as Gmail, Sheets, Airtable, GitHub, Slack, Calendar, Notion, and PagerDuty, covering accounting, auditing, batch editing, and synchronization. A task is considered passed only when all its fixed checks succeed.
Three dimensions, three different winners
Success rate (out of 30):
Oh My Pi: 17/30 (highest)
Claude Code: 16/30
Codex: 16/30
OpenCode: 14/30 (lowest)
Most counter‑intuitive: for seven tasks the pass/fail outcome depends entirely on which harness is used—same model, different wrapper changes the result.
Cost per successful task (estimated from API pricing):
OpenCode: $0.073 (cheapest)
Codex: $0.081
Oh My Pi: $0.103
Claude Code: $0.195 (most expensive, about 2.7 × the cheapest)
All four harnesses stay below $0.20 per successful task, but the author warns that DeepSeek’s price increase could break this cost pattern.
Median end‑to‑end time per task:
Claude Code: 122.7 s (fastest)
OpenCode: 129.7 s
Codex: 245.0 s
Oh My Pi: 272.4 s (slowest)
No all‑round winner
Comparing the three dimensions shows that the fastest is also the most expensive, and the highest success rate is the slowest. Claude Code is fastest but most costly; Oh My Pi has the highest success rate but is slowest; OpenCode is cheapest but has the lowest success rate. No harness dominates all three metrics.
Conclusion: the choice of harness changes the outcome of three tasks, nearly triples cost, and alters speed by 2.2×. Therefore, select the harness based on the metric you value most rather than a single overall ranking.
Composio - We ran DeepSeek V4 Flash through 4 agent harnesses (Claude Code, Codex, OpenCode, Oh My Pi) on 30 agentic tasks
https://x.com/composio/status/2085330847951970801Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
