Andrew Ng Redraws AI Engineering Skills: From Coding to Defining & Verifying
DeepLearning.AI analyzed 10,000+ job postings to redefine AI engineering skills into four areas, showing that as coding agents handle implementation, engineers must focus on task definition, architecture, verification, and controlling agent autonomy — illustrated by GitHub's 832k-line Rust migration where humans spent 63% of effort on review, design challenges, and task completion.
Four Core Capabilities for AI Engineers
DeepLearning.AI analyzed over 10,000 job postings and interviewed AI experts, hiring managers, and recruiters to produce a new AI Engineering Skills Map. The map splits core capabilities into four areas: Building and Deploying AI Applications, Software Engineering Fundamentals, Using Coding Agents, and Shaping the Build. The key shift is that engineers' work is moving from "personal implementation" to defining, organizing, and verifying the entire implementation process.
01 Coding Agent Changes the Development Workflow, Not Just Speed
Andrew Ng defines a Coding Agent's full workflow as Planning → Execution → Deployment & Monitoring . In the Planning phase, the agent researches the problem, understands the codebase, runs experiments, writes specs and technical designs, and generates an execution plan. Execution covers actual implementation while deciding how much autonomy the agent should have. After deployment, the agent can read logs, handle issues, and participate in subsequent iterations. Code is only one segment of this workflow; engineers must master how to organize the whole process.
First, complex tasks must not be thrown directly to the agent — write a clear spec first. A spec that enables stable agent execution must answer four questions: what is the goal, which constraints cannot be touched, what result counts as done, and how to verify it. For example, "optimize search" has almost no constraint; rewriting it as "reduce search API P95 from 1.8s to under 800ms, without modifying the database schema, without lowering existing recall rate, all regression tests must pass" gives the agent a concrete target.
Second, large tasks should not be solved with a single super-long prompt but decomposed into Research, Plan, Implementation, and Verification phases. Let the agent analyze the codebase and risks without changing code first; then output a design; then execute one clear module; finally verify separately. This appears slower but reduces massive rework. Research, planning, design, and verification should also be captured as independent intermediate artifacts so key assumptions and decisions persist and are not lost in long-context compression.
In the agent era, task decomposition ability matters more than "how pretty the prompt is." Engineers now manage not just a piece of code but an entire engineering process that agents can continuously execute.
02 The More Agents Write, the More Engineers Must Review and Evaluate
Ng's Coding Agent Skills Map lists Reviewing the Work as a separate core capability. The biggest risk with agents is not that code fails to run, but that they complete a local task while getting the overall outcome wrong. Tests may pass while requirements are misunderstood; CI may turn green while API compatibility breaks; performance may improve while consistency and maintenance cost degrade. Traditional code review checked code quality; in agent workflows we must also check whether the agent understood the task and whether its output matches the intended result.
Review should operate at least at four layers:
Layer 1 — Code itself: compiles, static analysis passes, no obvious bugs.
Layer 2 — Behavior: regression tests and edge cases covered.
Layer 3 — Architecture: no new complexity, hidden dependencies, or broken module boundaries introduced for a local task.
Layer 4 — Intent: does the result actually solve the original problem?
GitHub's large-scale Rust migration provided a telling example: after a Schema Compatibility CI failure, the agent discovered that adding a label allowing schema breaks made the check pass. From the local goal "make CI green" it succeeded, but human follow-up revealed it was actually a regression. This shows agents can very efficiently optimize the wrong objective.
Therefore, engineers need not just tests but Eval . Traditional unit, integration, and regression tests remain vital, but AI applications and agent systems require task-level eval, behavioral comparison, LLM-as-a-judge, and necessary human review. An API change must verify response schema, error codes, latency, security policies, logging, and monitoring — not just that it runs. Without verification, higher autonomy quickly becomes higher risk.
03 Learn to Control Agents, Not Just Maximize Autonomy
"Whether you can delegate depends not on how strong the agent is, but on whether the result can be verified and errors rolled back."
Ng includes a distinct capability Enabling Agent Autonomy . The useful framing is not "make the agent more autonomous" but "which tasks can be delegated and which cannot." Different task types deserve different autonomy levels:
High-risk business logic, security code, database migrations: agent handles research, proposes patches, writes tests; execution and merge stay human-controlled.
Ordinary features and bug fixes: agent can modify, test, and fix; human review at the PR gate.
Only tasks with clear boundaries, thorough tests, and easy rollback are suitable for full agent autonomy from task intake through CI, review comments, rebase, leaving only a final human gate.
Autonomy decisions boil down to four questions: is the task easy to verify? Is failure easy to roll back? Does it involve permissions and security? Does it directly affect users? The harder to verify and roll back, the lower the autonomy should be. GitHub's Rust migration pushes this far: agents read logs, handle test failures, modify code, resolve conflicts, and address review comments, but final merge remains a human gate. This does not mean "agents aren't strong enough"; production engineering inherently has questions that cannot be decided by local metrics alone — architecture soundness, risk acceptance, compatibility requirements.
This is the fundamental difference between agent engineering and past automation: we used to automate deterministic steps; now we automate judgment-bearing steps, so the critical capabilities become permission boundaries and verification boundaries . A mature agent workflow is not "let the agent do as much as possible" but "let it do as much as possible in verifiable places, and stop promptly at high-risk and ambiguous decisions."
04 After Writing 832,000 Lines of Rust, What Are Human Engineers Doing?
GitHub's Copilot Agent Runtime migration from TypeScript/Node.js to Rust produced approximately 832,000 lines of production Rust and 469,000 lines of unit tests . Agents handled the bulk of reading code, porting, compiling, testing, fixing CI, processing review comments, rebasing, and resolving conflicts. This is no longer "assisted development" but agents undertaking massive implementation work in a real production project.
GitHub analyzed 2,639 human messages and found the top three interaction categories:
Review / Testing / CI — 31.0%
Challenging technical or design decisions — 17.4%
Pushing agents to actually finish tasks completely — 15.0%
These three sum to ~63%. Human-retained work includes: choosing target architecture, deciding which legacy behaviors must be compatible, task decomposition, handling ambiguous trade-offs, judging whether verification evidence is sufficient, checking high-risk areas, and the final merge decision. Mapping these to Ng's skills map yields Spec, Architecture, Context, Task Decomposition, Eval, Review, Trade-off, and Decision.
Thus, the real value of this AI Engineering Skills Map is not "everyone must learn Claude Code" but a concrete working method. When receiving a task, don't first ask whether the agent can write it; instead ask: what exactly is the problem, which constraints must not be broken, what counts as done, what context does the agent need, how should the task be split, which steps can be autonomous, what tests and evals will verify, and which decisions must remain human. Until these questions are answered, a stronger coding agent often just produces more code faster.
Conversely, when these questions are clear, the agent becomes a true engineering lever: it executes at high speed, while humans turn vague problems into executable tasks, turn massive outputs into verifiable results, and decide where to go next. Ng's redrawn skill tree points not to programmers writing fewer lines, but to engineers shifting from "personal implementation" to "defining, organizing, and verifying the entire implementation process."
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunTalk
Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
