Step 5 Preview Outperforms DeepSeek V4 Pro in Coding Tasks at Fraction of Opus 5 Cost

The article benchmarks StepFun's Step 5 Preview against DeepSeek V4 Pro across three coding challenges—2D animation, token bucket demo, and CSV analysis workbench—revealing Step 5 Preview delivers richer details and better test coverage despite longer generation times, with per-task cost at just 35% of GLM-5.3 and 12.5% of Claude Opus 5.

JavaGuide
JavaGuide
JavaGuide
Step 5 Preview Outperforms DeepSeek V4 Pro in Coding Tasks at Fraction of Opus 5 Cost

Introduction

The author evaluates StepFun's newly released Step 5 Preview model through three practical coding tasks, comparing it side-by-side with DeepSeek V4 Pro. Each task uses identical prompts but different toolchains: Step 5 Preview via WorkBuddy with browser-based verification, DeepSeek V4 Pro via local DeepSeek Harness with syntax checks only. A third-party model (GLM) provides blind evaluations.

Step 5 Preview Model Overview

Step 5 Preview is positioned as "Advancing the Pareto Frontier" — a sparse Mixture-of-Experts (MoE) model with 600B total parameters, 27B activated per token, 1M token context window, and vision input support. Official benchmarks show Step 5 Preview (High) scoring 67.7 on DeepSWE v1.1, 80.5 on ProgramBench, and 49.0 on StepCodeBench (avg@4), surpassing Kimi K3 (43.9) and GLM-5.3 (40.2), though trailing Opus 5 (63.9). Cost per task is reported at 35% of GLM-5.3/Kimi K3 and 12.5% of Claude Opus 5.

Step 5 Preview official introduction showing model positioning and architecture
Step 5 Preview official introduction showing model positioning and architecture
Official eight-benchmark comparison preserving model versions and thinking settings
Official eight-benchmark comparison preserving model versions and thinking settings
StepCodeBench top six task types comparison with overall avg@4 and per-type trends
StepCodeBench top six task types comparison with overall avg@4 and per-type trends
Intelligence vs. per-task cost scatter plot with log-scaled dollar cost on x-axis and Artificial Analysis Intelligence Index on y-axis
Intelligence vs. per-task cost scatter plot with log-scaled dollar cost on x-axis and Artificial Analysis Intelligence Index on y-axis

WorkBuddy Integration Setup

To call Step 5 Preview, the author creates an API key on the StepFun open platform ( https://platform.stepfun.com/), then configures a custom model in WorkBuddy with:

API endpoint: https://api.stepfun.com/v1/chat/completions (Step Plan uses https://api.stepfun.com/step_plan)

API key from the platform

Model name: step-5-preview Advanced: enable tool calling; use provider defaults for input/output length

WorkBuddy custom model configuration showing endpoint, model name, and tool calling option
WorkBuddy custom model configuration showing endpoint, model name, and tool calling option

Case 1: 2D Animation — Sun Wukong Pilots Plane, Pelican Passenger

Step 5 Preview Output

Completed in 30 minutes 3 seconds . Uses Canvas 2D rendering. The animation depicts Sun Wukong with golden headband, monkey face, and tiger-striped clothing in the pilot seat; the pelican in the rear seat has a long beak, throat pouch, flight cap, goggles, and red scarf. Control stick, wings, tail, background mountains, and birds are all drawn. A single timeline drives all actions. Verification via headless Chrome and pixel checks caught two issues: premature phase-label change and reversed landing-gear animation, both fixed.

Step 5 Preview generated airplane animation with Sun Wukong in front, pelican in rear
Step 5 Preview generated airplane animation with Sun Wukong in front, pelican in rear

DeepSeek V4 Pro Output

Completed in 9 minutes 49 seconds via local Harness. Also uses Canvas and a unified animation clock. The screenshot shows runway and wheels still visible; Sun Wukong identified by headband, monkey face, and golden staff; pelican retains beak and throat pouch; two dark cockpit areas distinguish front/rear seats. However, the harness only performed node --check syntax validation — no actual browser preview or real-device acceptance test.

DeepSeek V4 Pro High takeoff scene with Sun Wukong and pelican in front and rear seats
DeepSeek V4 Pro High takeoff scene with Sun Wukong and pelican in front and rear seats

Comparison

Step 5 Preview shows richer character accessories and environmental details. DeepSeek uses simpler shapes but satisfies role/seat requirements. The two screenshots capture different flight phases (cruise vs. takeoff), so direct visual comparison is limited. On the "trigger turbulence" button during pause, Step 5 Preview logs the turbulence window while DeepSeek ignores the click. Dynamic behavior and button states need identical interactive testing; static images cannot confirm full compliance. Time difference (30m vs 9m) reflects toolchain and verification scope differences, not raw model speed alone.

Case 2: Token Bucket Rate Limiting Interactive Demo

Task: Build a Vite + TypeScript Chinese interactive demo. Default capacity 5, refill 2 tokens/second, time advanced only by button clicks. Required sequence: request 3, request 3, advance 1s, request 4, advance 5s, request 5. Expected final state: 0 tokens, virtual time 6s, 3 allowed, 1 rejected. Rejection must not deduct tokens; refill must not exceed capacity; reset must clear logs and stats.

Step 5 Preview Output

Time: 29 minutes 5 seconds . UI: left side shows bucket level with liquid and dots; right side groups config and action buttons; bottom has independent scrolling log area. A six-step guide highlights the next pending action. Screenshot shows after two requests: tokens 5→2, second request rejected (tokens stay 2), time 0s, allowed 1, rejected 1, progress "In progress 3/6" with hint "Next: advance 1s".

Step 5 Preview token bucket demo after two requests, 2 tokens remaining, next step advance 1s
Step 5 Preview token bucket demo after two requests, 2 tokens remaining, next step advance 1s

Delivery report: 37 test cases passed , type-check and build successful, Chromium verification of six-step demo, invalid input, reset, and 1440×900 layout. Fixed a long-log layout overflow by constraining log scrolling to its own region.

Step 5 Preview execution and delivery notes showing 29m5s, 37 tests, browser checks, and layout fix
Step 5 Preview execution and delivery notes showing 29m5s, 37 tests, browser checks, and layout fix

DeepSeek V4 Pro Output

Time: 11 minutes 6 seconds , 45 tool calls . Three-column layout: bucket state, config/actions, operation log. Color-coded log entries for allow, reject, time advance. Screenshot shows after step 3 (advance 1s): tokens 4, virtual time 1s, log consistent, stats allow 1 reject 1. Progress shows "Completed 3/6 steps" (different meaning from Step 5 Preview's "In progress 3/6").

DeepSeek token bucket demo after step 3, 4 tokens, virtual time 1s
DeepSeek token bucket demo after step 3, 4 tokens, virtual time 1s

Delivery report: 20 automated tests passed (10 algorithm, 10 controller), type-check and build pass, 36 browser checks via local Chrome. Notes: layout checks via DOM size and scroll assertions, no pixel visual diff, no mobile adaptation; desktop-only per requirements.

DeepSeek execution and delivery notes showing 11m6s, 20 tests, 36 browser checks
DeepSeek execution and delivery notes showing 11m6s, 20 tests, 36 browser checks

Blind Evaluation (Third-Party Model)

Functional verdict: tie. Both pass identical six-step demo, reset, invalid input checks, and re-run each other's tests. Differences lie in pedagogical expression and engineering details. Step 5 Preview explicitly indicates how many tokens are dropped when refill hits capacity, and provides a six-step checklist. Reviewer notes Step 5 Preview restricts manual actions during demo, covers every intermediate state and log-cap tests — more complete. DeepSeek's strength: concise implementation meeting explicit requirements with less code.

Case 2 blind evaluation: S=Step 5 Preview, D=DeepSeek V4 Pro
Case 2 blind evaluation: S=Step 5 Preview, D=DeepSeek V4 Pro

Case 3: CSV Order Analysis Workbench from Scratch

Scope: import order CSV, detect errors, filter/analyze, export current results. Both start in empty directories with Vite + TypeScript, same 18 valid orders and error sample, no backend. Tricky data: customer names contain commas, double quotes, newlines; payment/refund/cancel orders counted separately; amounts in integer cents; error file must not overwrite imported data. Pagination 10 rows/page, but export includes all filtered records and must be re-importable.

Step 5 Preview Output

UI: file ops at top, then filters, four metric cards, two charts, order table. Chinese column headers, date range inclusive, export shows record count toast. Time: 1 hour 26 minutes . 56 automated tests passed , type-check and build pass, system Chrome verification of import, pagination, combined filtering, error recovery, chart linkage, export of 12 records and re-import.

Step 5 Preview CSV workbench filtering Sep 1-3, showing 9 records and export success toast
Step 5 Preview CSV workbench filtering Sep 1-3, showing 9 records and export success toast
Step 5 Preview execution and delivery notes with 56 tests, browser checks, and follow-up confirmations
Step 5 Preview execution and delivery notes with 56 tests, browser checks, and follow-up confirmations

DeepSeek V4 Pro Output

Similar layout: filter area, metric cards, charts, order table. Daily net income bars show values directly, positive/negative vs zero axis easy to read. Table headers remain English ( date, customer, amount, etc.) despite Chinese UI requirement. Screenshot shows no date filter applied: all 18 records, payment ¥1,438, refund ¥458.50, net ¥979.50 — matches ground truth. Six daily net values: 199, 348, -141, 218, 178, 177.50; chart labels correspond.

DeepSeek CSV workbench showing all 18 orders and daily net income with value labels
DeepSeek CSV workbench showing all 18 orders and daily net income with value labels

Time: 13 minutes 25 seconds , 72 tool calls . 49 unit tests + 34 browser end-to-end checks passed , type-check and build pass. Coverage includes error import preserving old data, 10/8 pagination, cross-page export of 12 records, re-import consistency. Notes: model cannot view screenshots, so no pixel visual diff; browser tests rely on DOM assertions and visibility; local Chrome required, documented in README.

DeepSeek execution and delivery notes showing 13m25s, 49 unit tests, 34 browser checks
DeepSeek execution and delivery notes showing 13m25s, 49 unit tests, 34 browser checks

Blind Evaluation (GLM)

Anonymized: A=DeepSeek, B=Step 5 Preview. GLM first verified CSV byte-for-byte identity, then ran 32 core logic assertions and 80 browser operations on both, plus each project's own tests/builds. Core logic: 32/32 both. Browser checks: 79/80 both. Key behaviors (payment/refund stats, cross-page export, error import retention) all pass.

Issues found: DeepSeek resets filters on new import but leaves stale keyword "Zhang San" in search box; Step 5 Preview only shows this briefly when import triggered while input focused, clears on blur. Step 5 Preview advantages: duplicate ID error cites conflicting row; empty-state chart shows "no data"; bar chart keyboard operable; 1440×900 fits without page scroll. DeepSeek requires page scroll but buttons visible, table independently scrollable — still meets layout spec. DeepSeek delivers reproducible end-to-end scripts (34/34 re-run pass); Step 5 Preview has more unit tests but no equally reproducible browser scripts. Test count ≠ verification reproducibility.

Case 3 GLM blind evaluation: A=DeepSeek, B=Step 5 Preview, with acceptance results and state sync issues
Case 3 GLM blind evaluation: A=DeepSeek, B=Step 5 Preview, with acceptance results and state sync issues

Blind evaluation conclusion: Step 5 Preview slightly better, gap very small . Both complete main data flows; differences center on state consistency, error localization, empty-state messaging, and whether browser verification can be reproduced with the project.

Conclusion

Across three cases, Step 5 Preview delivers overall better results: Case 1 richer character/environment details; Case 2 six-step guide embedded in page; Case 3 clearer Chinese headers and filter/stat/export explanations. Trade-off: longer generation times — DeepSeek 9m49s, 11m6s, 13m25s vs Step 5 Preview 30m3s, 29m5s, 1h26m. Cost advantage remains: per-task cost 35% of GLM-5.3/Kimi K3, 12.5% of Claude Opus 5 — high value for money, worth trying.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Mixture of Expertstoken bucketcost efficiencyWorkBuddyDeepSeek V4 ProAI coding benchmarkStep 5 PreviewCSV analysis
JavaGuide
Written by

JavaGuide

Backend tech guide and AI engineering practice covering fundamentals, databases, distributed systems, high concurrency, system design, plus AI agents and large-model engineering.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.