Step 5 Preview Tested: AI Agent Builds 3D Frontend & Deploys Full Site Autonomously
The author tests Step 5 Preview by prompting it to create complex three.js 3D visualizations and build a complete showcase website for an open-source agent skills repository, including design proposals, code generation, headless browser verification, git operations, and deployment to GitHub Pages, demonstrating autonomous engineering capabilities at a fraction of the cost of top closed models.
3D Earth Visualization Test
The author accessed Step 5 Preview via Claude Code using the prompt create a highly detailed 3d earth using three.js. The model produced a scene featuring terrain relief, ocean specular highlights, terminator atmospheric scattering, atmospheric glow, night city lights, and a starfield. Interactions included inertia-damped drag rotation, scroll zoom, layer toggles, sun-angle switching, dark/light theme toggle, and five preset camera views. The author judged the result superior to a reference demo shown earlier.
3D Black Hole Visualization Test
With the prompt create a highly detailed 3d black hole using three.js, Step 5 Preview generated a physically detailed black hole simulation. The author noted the physics engine design was sophisticated. The interface provided 17 sliders, 5 presets, inertia-smoothed dragging, and keyboard controls.
Real-World Task: Building a Showcase Website for Agent Skills
The author tasked the model with creating a display and retrieval site for their open-source repository tjxj/z-skills (17 Agent Skills). The model first used WebFetch to inspect the GitHub repo, extracted all skill names, functions, and trigger words, and organized them into five practical categories while confirming a unified npx installation command. Step 5 Preview then paused to propose three distinct visual designs, even drawing ASCII wireframes for layout. It recommended a Terminal style, but the author chose an "Aurora Glass" theme. After writing the HTML, the model autonomously searched for an available environment, launched Chrome Headless, captured full-page and above-the-fold screenshots, and used its multimodal vision to verify visual quality like a professional tester.
Autonomous Verification and Deployment
Upon confirming zero visual defects, the model offered several deployment options. The author selected "Deploy to GitHub Pages." The model ran git status, detected untracked temporary directories and existing design docs, staged only docs/index.html, pushed the commit, and enabled Pages via the GitHub API.
Iterative Refinement and Final Assessment
The author found the Aurora Glass style too flashy and requested a return to the Terminal dark theme. The model noticed that secondary text on a deep-black background appeared gray and low-contrast, so it proactively brightened the gray scale for sharper readability. Further iterations added a dark/light mode toggle, a cheatsheet, a session demo, and a collapsible FAQ.
The author concludes that Step 5 Preview demonstrates complete autonomous delivery at an engineer-level agent standard: understanding a real open-source project, communicating design proposals, writing code, self-verifying with a headless browser, precise Git commits, production deployment, and handling picky feedback with self-debugging.
Cost-Performance Comparison
Previously, developers faced a dilemma: top-tier closed models offer superior comprehension and finishing ability but incur high costs for long-horizon tasks; cheaper lightweight models handle small functions but fail at complex multi-file repository understanding, autonomous headless browser self-checks, deep engineering logic fixes, and multi-version deployment coordination. Step 5 Preview currently ranks in the global top three open models on the AA comprehensive intelligence leaderboard, with intelligence comparable to K3 and GLM 5.3, while single-task cost is only 35% of theirs, offering strong value. The author suggests hands-on testing with typical tasks to form a personal judgment.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Old Zhang's AI Learning
AI practitioner specializing in large-model evaluation and on-premise deployment, agents, AI programming, Vibe Coding, general AI, and broader tech trends, with daily original technical articles.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
