Step 5 Preview Outperforms DeepSeek V4 Pro in Coding Tasks at Fraction of Opus 5 Cost
The article benchmarks StepFun's Step 5 Preview against DeepSeek V4 Pro across three coding challenges—2D animation, token bucket demo, and CSV analysis workbench—revealing Step 5 Preview delivers richer details and better test coverage despite longer generation times, with per-task cost at just 35% of GLM-5.3 and 12.5% of Claude Opus 5.
Introduction
The author evaluates StepFun's newly released Step 5 Preview model through three practical coding tasks, comparing it side-by-side with DeepSeek V4 Pro. Each task uses identical prompts but different toolchains: Step 5 Preview via WorkBuddy with browser-based verification, DeepSeek V4 Pro via local DeepSeek Harness with syntax checks only. A third-party model (GLM) provides blind evaluations.
Step 5 Preview Model Overview
Step 5 Preview is positioned as "Advancing the Pareto Frontier" — a sparse Mixture-of-Experts (MoE) model with 600B total parameters, 27B activated per token, 1M token context window, and vision input support. Official benchmarks show Step 5 Preview (High) scoring 67.7 on DeepSWE v1.1, 80.5 on ProgramBench, and 49.0 on StepCodeBench (avg@4), surpassing Kimi K3 (43.9) and GLM-5.3 (40.2), though trailing Opus 5 (63.9). Cost per task is reported at 35% of GLM-5.3/Kimi K3 and 12.5% of Claude Opus 5.
WorkBuddy Integration Setup
To call Step 5 Preview, the author creates an API key on the StepFun open platform ( https://platform.stepfun.com/), then configures a custom model in WorkBuddy with:
API endpoint: https://api.stepfun.com/v1/chat/completions (Step Plan uses https://api.stepfun.com/step_plan)
API key from the platform
Model name: step-5-preview Advanced: enable tool calling; use provider defaults for input/output length
Case 1: 2D Animation — Sun Wukong Pilots Plane, Pelican Passenger
Step 5 Preview Output
Completed in 30 minutes 3 seconds . Uses Canvas 2D rendering. The animation depicts Sun Wukong with golden headband, monkey face, and tiger-striped clothing in the pilot seat; the pelican in the rear seat has a long beak, throat pouch, flight cap, goggles, and red scarf. Control stick, wings, tail, background mountains, and birds are all drawn. A single timeline drives all actions. Verification via headless Chrome and pixel checks caught two issues: premature phase-label change and reversed landing-gear animation, both fixed.
DeepSeek V4 Pro Output
Completed in 9 minutes 49 seconds via local Harness. Also uses Canvas and a unified animation clock. The screenshot shows runway and wheels still visible; Sun Wukong identified by headband, monkey face, and golden staff; pelican retains beak and throat pouch; two dark cockpit areas distinguish front/rear seats. However, the harness only performed node --check syntax validation — no actual browser preview or real-device acceptance test.
Comparison
Step 5 Preview shows richer character accessories and environmental details. DeepSeek uses simpler shapes but satisfies role/seat requirements. The two screenshots capture different flight phases (cruise vs. takeoff), so direct visual comparison is limited. On the "trigger turbulence" button during pause, Step 5 Preview logs the turbulence window while DeepSeek ignores the click. Dynamic behavior and button states need identical interactive testing; static images cannot confirm full compliance. Time difference (30m vs 9m) reflects toolchain and verification scope differences, not raw model speed alone.
Case 2: Token Bucket Rate Limiting Interactive Demo
Task: Build a Vite + TypeScript Chinese interactive demo. Default capacity 5, refill 2 tokens/second, time advanced only by button clicks. Required sequence: request 3, request 3, advance 1s, request 4, advance 5s, request 5. Expected final state: 0 tokens, virtual time 6s, 3 allowed, 1 rejected. Rejection must not deduct tokens; refill must not exceed capacity; reset must clear logs and stats.
Step 5 Preview Output
Time: 29 minutes 5 seconds . UI: left side shows bucket level with liquid and dots; right side groups config and action buttons; bottom has independent scrolling log area. A six-step guide highlights the next pending action. Screenshot shows after two requests: tokens 5→2, second request rejected (tokens stay 2), time 0s, allowed 1, rejected 1, progress "In progress 3/6" with hint "Next: advance 1s".
Delivery report: 37 test cases passed , type-check and build successful, Chromium verification of six-step demo, invalid input, reset, and 1440×900 layout. Fixed a long-log layout overflow by constraining log scrolling to its own region.
DeepSeek V4 Pro Output
Time: 11 minutes 6 seconds , 45 tool calls . Three-column layout: bucket state, config/actions, operation log. Color-coded log entries for allow, reject, time advance. Screenshot shows after step 3 (advance 1s): tokens 4, virtual time 1s, log consistent, stats allow 1 reject 1. Progress shows "Completed 3/6 steps" (different meaning from Step 5 Preview's "In progress 3/6").
Delivery report: 20 automated tests passed (10 algorithm, 10 controller), type-check and build pass, 36 browser checks via local Chrome. Notes: layout checks via DOM size and scroll assertions, no pixel visual diff, no mobile adaptation; desktop-only per requirements.
Blind Evaluation (Third-Party Model)
Functional verdict: tie. Both pass identical six-step demo, reset, invalid input checks, and re-run each other's tests. Differences lie in pedagogical expression and engineering details. Step 5 Preview explicitly indicates how many tokens are dropped when refill hits capacity, and provides a six-step checklist. Reviewer notes Step 5 Preview restricts manual actions during demo, covers every intermediate state and log-cap tests — more complete. DeepSeek's strength: concise implementation meeting explicit requirements with less code.
Case 3: CSV Order Analysis Workbench from Scratch
Scope: import order CSV, detect errors, filter/analyze, export current results. Both start in empty directories with Vite + TypeScript, same 18 valid orders and error sample, no backend. Tricky data: customer names contain commas, double quotes, newlines; payment/refund/cancel orders counted separately; amounts in integer cents; error file must not overwrite imported data. Pagination 10 rows/page, but export includes all filtered records and must be re-importable.
Step 5 Preview Output
UI: file ops at top, then filters, four metric cards, two charts, order table. Chinese column headers, date range inclusive, export shows record count toast. Time: 1 hour 26 minutes . 56 automated tests passed , type-check and build pass, system Chrome verification of import, pagination, combined filtering, error recovery, chart linkage, export of 12 records and re-import.
DeepSeek V4 Pro Output
Similar layout: filter area, metric cards, charts, order table. Daily net income bars show values directly, positive/negative vs zero axis easy to read. Table headers remain English ( date, customer, amount, etc.) despite Chinese UI requirement. Screenshot shows no date filter applied: all 18 records, payment ¥1,438, refund ¥458.50, net ¥979.50 — matches ground truth. Six daily net values: 199, 348, -141, 218, 178, 177.50; chart labels correspond.
Time: 13 minutes 25 seconds , 72 tool calls . 49 unit tests + 34 browser end-to-end checks passed , type-check and build pass. Coverage includes error import preserving old data, 10/8 pagination, cross-page export of 12 records, re-import consistency. Notes: model cannot view screenshots, so no pixel visual diff; browser tests rely on DOM assertions and visibility; local Chrome required, documented in README.
Blind Evaluation (GLM)
Anonymized: A=DeepSeek, B=Step 5 Preview. GLM first verified CSV byte-for-byte identity, then ran 32 core logic assertions and 80 browser operations on both, plus each project's own tests/builds. Core logic: 32/32 both. Browser checks: 79/80 both. Key behaviors (payment/refund stats, cross-page export, error import retention) all pass.
Issues found: DeepSeek resets filters on new import but leaves stale keyword "Zhang San" in search box; Step 5 Preview only shows this briefly when import triggered while input focused, clears on blur. Step 5 Preview advantages: duplicate ID error cites conflicting row; empty-state chart shows "no data"; bar chart keyboard operable; 1440×900 fits without page scroll. DeepSeek requires page scroll but buttons visible, table independently scrollable — still meets layout spec. DeepSeek delivers reproducible end-to-end scripts (34/34 re-run pass); Step 5 Preview has more unit tests but no equally reproducible browser scripts. Test count ≠ verification reproducibility.
Blind evaluation conclusion: Step 5 Preview slightly better, gap very small . Both complete main data flows; differences center on state consistency, error localization, empty-state messaging, and whether browser verification can be reproduced with the project.
Conclusion
Across three cases, Step 5 Preview delivers overall better results: Case 1 richer character/environment details; Case 2 six-step guide embedded in page; Case 3 clearer Chinese headers and filter/stat/export explanations. Trade-off: longer generation times — DeepSeek 9m49s, 11m6s, 13m25s vs Step 5 Preview 30m3s, 29m5s, 1h26m. Cost advantage remains: per-task cost 35% of GLM-5.3/Kimi K3, 12.5% of Claude Opus 5 — high value for money, worth trying.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
JavaGuide
Backend tech guide and AI engineering practice covering fundamentals, databases, distributed systems, high concurrency, system design, plus AI agents and large-model engineering.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
