Plan Mode Evolved: The Verifiable Work Loop AI Agents Actually Need

The article argues that Plan Mode isn't obsolete but should shift from static pre‑approval documents to a dynamic, evidence‑updated work loop where agents explore, act, verify, and escalate high‑impact decisions to humans, illustrated with a sync‑to‑async export refactoring example.

Architect
Architect
Architect
Plan Mode Evolved: The Verifiable Work Loop AI Agents Actually Need

Introduction

On September 24, Ayman Nadeem published "Plan mode is dead," reflecting on a product experiment at Nuanced. The title sounds like a product verdict, but a close read shows Ayman is discussing a specific R&D workflow: generate a full plan, get human approval, then hand it to an agent for execution.

That workflow used to be valuable. However, as agents become capable of reading code, trying approaches, running tests, and adjusting based on results, the more complete a plan is written before execution starts, the more likely it becomes an outdated assumption that must be maintained when new facts appear.

Planning is still needed; the question is whether it should remain centered on a single long document.

Official product docs show Plan Mode hasn't disappeared from agent tools. Claude Code still offers a plan permission mode where the agent can read files, run read‑only commands, and organize a plan but cannot modify source files. The same permission system includes an auto mode that lets the model pre‑approve a subset of tool calls. Google's Gemini Code Assist Agent Mode also retains explicit planning: complex tasks can first show a high‑level plan for user comment, edit, and approval of both the plan and tool calls.

Viewing these two product designs together sharpens the question: when agents can explore and act on their own, where do humans need to see key decisions, and which actions still require intervention before they happen?

We used to worry that agents wouldn't think first; now the more common frustration is that they write a very complete plan but keep encountering new facts during execution. The longer the plan, the harder it is for humans to distinguish confirmed decisions from guesses the model made based on current context.

The author cares about what responsibility planning actually carries in an agent workflow. The essence a plan must retain is goals, invariants, unknowns, verification evidence, and forks that require human decisions. Planning therefore becomes more like a verifiable work loop than a pre‑start document.

First Look at a Real Development Process

Take a common R&D task: converting a synchronous export interface into a background job.

The old way: have the agent write a complete proposal, humans annotate the document, approve it, then let the agent change code. A process better suited to agents is to first clarify the goal and boundaries, then let the agent gather evidence with low‑risk actions; when evidence changes the judgment, update the current state, and only ask humans to decide at high‑impact forks.

Agent R&D task verifiable work loop
Agent R&D task verifiable work loop

Figure 1 | A verifiable work loop for an agent R&D task.

In this chain, planning hasn't vanished; it's been distributed to different positions. Goals and constraints stay in the task state, exploration actions produce new facts, permission control guards high‑impact operations, tests and logs prove results, and humans intervene only at forks that carry product or architectural consequences.

Phase Breakdown

Define Outcome — Agent rewrites requirements into goals, constraints, and acceptance criteria. Artifacts: Goal, acceptance criteria, scope. Human looks when goal involves public commitment or cross‑team agreement.

Explore System — Agent reads code, traces call chains, runs read‑only checks. Artifacts: Confirmed facts, unknowns, impact scope. Human looks when original assumptions are overturned by code facts.

Small Actions — Agent adds tests, makes local changes, runs rollback‑safe experiments. Artifacts: Diffs, command output, test results. Human looks when action will touch data, permissions, or public APIs.

Update State — Agent adjusts next steps and stop conditions based on evidence. Artifacts: Decision log, remaining risks. Human looks when evidence conflicts or several rounds show no progress.

Handoff — Agent organizes changes, verifies results, lists open issues. Artifacts: Reviewable work site. Human looks before merge, release, or responsibility transfer.

Separate Planning from Plan

Ayman's article makes a distinction worth keeping: planning and plan are not the same thing.

The former is the continuous process of understanding, judging, trying approaches, and absorbing new facts; the latter is usually a written plan that preserves a subset of conclusions and hands them to subsequent execution.

Nuanced's original problem hasn't expired: when machines modify software faster than humans can check changes, how do people maintain system understanding?

The problem lies in the implementation. If you fix the whole process as:

Conversation → Disambiguate → Generate Spec → Review → Approve → Implement → Code Review

it implies a premise: understanding can be completed before acting, and execution merely lands what's already decided.

Software development is rarely that neat. You read a call chain and discover the original assumption doesn't hold; you add a test and find the interface's error semantics can't be changed; you run real data and learn the "simple migration" affects downstream clients.

Execution is not a downstream factory of planning. Execution itself generates new information, and new information changes planning.

Therefore, the outdated part is forcing thinking into a phase that can only happen before work starts.

Why Plan Mode Worked Before

Plan Mode used to be useful for straightforward reasons. It did two jobs at once.

One job faced the agent: writing goals, file paths, module boundaries, and execution order clearly reduced the chance the model would fill in wrong assumptions from local code. Early models weren't good at exploring large codebases and tended to start editing files as soon as they saw a requirement, so precise instructions were valuable.

The other job faced humans: before large amounts of code were generated, it helped developers confirm "is this what I want?" and build a preliminary understanding of the system.

Now, the necessity of the first job is decreasing because models increasingly read code, form hypotheses, run tools, and correct plans on their own; the necessity of the second job is increasing because code generation is faster, making it harder for humans to keep up with decisions, assumptions, and impact scope.

The problem is that these two jobs have long been stuffed into the same plan document. The model doesn't necessarily need a long instruction, and humans don't necessarily build a reliable mental model by reading a long instruction.

Hence the early workflow naturally summarized as:

Plan First → Human Approve → Agent Execute
Approval plan vs action loop comparison
Approval plan vs action loop comparison

Figure 2 | Difference between two work shapes. The plan still exists, but shifts from a one‑time approval text to a work state updated with evidence.

This chain front‑loads decisions, reducing the chance the model acts continuously in a wrong direction. That value still exists today. OpenAI's internal Codex practice suggests large changes first form an implementation plan in Ask Mode, then bring it into Code Mode. Anthropic's current permission docs still provide plan mode, allowing agents to read files and run read‑only commands but not modify source. Google's Gemini Code Assist Agent Mode also keeps plans and tool use as collaborative nodes that can be commented, edited, and approved.

Hacker News discussion didn't converge on a single answer. Some treat Plan Mode as a read‑only exploration and review checkpoint; others think it's just a prompt wrapper and saying "plan first" in chat is enough. The divergence itself shows people aren't using the same Plan Mode.

So "Plan Mode is dead" is not an industry fact; it's a product author's retrospective judgment based on their own workflow.

Why Plans Started Getting Heavy

As model capabilities improved, plans began to carry an awkward burden.

They originally helped agents think through how to do things; now models increasingly read code, propose hypotheses, run small experiments, and correct plans themselves. They also originally helped humans understand the system, but AI‑generated long texts often make it harder for people to focus.

More information does not equal more clarity.

Plans get heavy often because of excessive detail: file paths, implementation steps, edge conditions, pseudocode, background explanations all go in, yet readers struggle to tell which are confirmed facts and which are guesses based on current context.

The longer the plan, the more assumptions hide inside it.

Nuanced even tried a Spec Tour to guide users through highlights of the spec document. That attempt is telling: if a complete spec still needs a guided tour to be understood, the product may just be adding another layer of text to watch rather than a clearer decision chain.

Worse, after plan approval, new evidence from execution has no natural place to be written back. The agent then faces two paths:

Faithfully execute a plan already overturned by new facts;

Pause, push the problem upstream, and rerun the whole approval process.

The first solidifies wrong assumptions into code; the second makes the whole workflow rigid.

The question can be reframed more practically: Which parts of the plan need to be persisted, and which are just temporary products of the current reasoning round?

How an R&D Task Passes Through This Loop

Again, look at "convert synchronous export to background task."

Define Outcome First, Not Files

Step one doesn't need to list files to change immediately. First write the outcome clearly: large file exports no longer occupy synchronous requests, while preserving old client error‑code semantics, retries must not create duplicate tasks, and existing task‑status interfaces must not be silently broken.

This aligns with OpenAI's current Goal use‑case direction: let Codex work across multiple turns around a persistent goal, using verifiable stop conditions to judge completion. Tasks should specify outcomes, constraints, verification methods, and blocking conditions — better than "implement an async export module" which only states an action, not what completion means.

Read Call Chains Before Deciding Approach

The agent can first read the gateway, export service, task table, and client contract tests, run read‑only checks, and draw a minimal call chain: where requests enter, where files are generated, where task status is stored, who interprets error codes.

The most important artifacts of this phase are three kinds of facts: confirmed constraints, remaining unknowns, and the next hypothesis most worth verifying.

After researching the repo, you might discover:

The gateway currently streams files directly to clients, and old clients depend on this behavior;

The retry mechanism lacks idempotency keys, so duplicate submissions may generate two files;

The task‑status interface already exists but is used by another module;

Failure error codes are treated by the frontend as retry signals.

At this point, the original "add task table, integrate queue, add status interface" plan is already incomplete. The problem isn't that the agent failed to execute the plan; repo facts changed the plan's premises.

Use Small Actions to Get Evidence

Next steps can be: add an old‑client contract test, experiment with an idempotency key for task submission, or just run the gateway‑to‑task‑table call chain locally.

These actions don't immediately change public APIs but answer key questions: who exactly depends on the old behavior, can existing fields express idempotent state, can the task‑status interface really be reused.

If evidence supports the original plan, the agent proceeds; if evidence overturns assumptions, update the work state. At that point, recording the change as a few checkable items is more valuable than regenerating a multi‑thousand‑word plan.

goal: Convert large file export from synchronous request to background task

invariants:
- Old client error code semantics must not change silently
- Retries must not create duplicate tasks
- Existing task status interface must remain compatible

unknowns:
- Whether gateway timeout config allows keeping old interface for a while
- Whether existing task table fields can express idempotent state

next_action:
- Check export call chain, task table, and client contract tests

evidence_needed:
- Retry scenario tests pass
- Old client contract tests pass
- Large file export no longer occupies synchronous request

human_checkpoint:
- Confirm how long old and new interfaces run in parallel

This record doesn't look like a complete spec, yet it's closer to the work state agents and humans need to share. It tells the agent the current goal, what must not be broken, and what facts are missing; it also tells humans which question they'll need to judge at the next intervention.

Leave High‑Impact Decisions to Humans

If the gateway can't stably support the old interface in parallel, the agent can keep building compatibility layers and tests, but "how long to keep the old interface" becomes a product and architecture decision that must stop at a human checkpoint.

If the final decision is to keep both interfaces, subsequent implementation must address compatibility period, traffic switching, and rollback paths; if the decision is to upgrade clients directly, task boundaries and acceptance criteria must also update. These records can live in Markdown, Issues, a task database, or the agent's own structured state. The key is that every change points to a new piece of evidence and leaves the next step and remaining risks.

From Approval Plan to Action Loop

Ayman's alternative path can be summarized as:

Understand → Act → Check → Clarify → Adjust → Act Again

The most valuable part of this loop is putting "execution feedback" back into the planning process.

To ground this loop in agent engineering, the author prefers recording five state categories:

Intent — What user behavior or system property this run aims to change.

Constraints — Which APIs, data, permissions, and performance conditions must not be broken.

Actions — What to read, modify, or experiment with next.

Evidence — What tests, logs, screenshots, diffs, or external receipts show.

Forks — Which results let the agent continue, which require human decision.

Seen this way, agents don't need to write the future completely before executing a plan; instead, after each action they should use evidence to update their understanding of the problem.

At runtime, this loop must have at least three exits . Local, rollback‑safe, clearly verifiable actions can continue; forks involving public APIs, data, permissions, or business commitments go to humans; when tests can't run, dependencies are unreadable, scope exceeds authorization boundaries, or several rounds yield no new evidence, it's better to stop than to mask uncertainty with more text.

This loop also connects with several threads discussed over recent months: Goal solves what long tasks aim for and when they can stop; Harness solves what environment agents act in, what tools they can use, and how they get feedback; Work Site solves what state must persist for the next human or agent to pick up; Spec and Rules solve how intent, boundaries, and acceptance criteria are made explicit. Each layer carries distinct responsibilities and cannot replace each other.

One agent job can be split into several responsibility‑clear layers: Goal owns direction, Rules own boundaries, Tools own actions, Verification owns evidence, Runtime owns lifecycle, Humans own high‑impact decisions.

Agent work loop architecture layers
Agent work loop architecture layers

Figure 3 | Architecture layers of the agent work loop; Plan Mode is just one entry or pause point.

In this diagram, Plan Mode is merely an entry or pause point. What sustains long tasks is the connection between goals, state, tools, evidence, and human decisions.

Architecture Pieces and Their Roles

Goal — What result the task must achieve, when it can stop. Without it: Agent keeps doing local optimizations but never has a completion judgment.

Work State — What is currently known, what's missing, what's next. Without it: Every round re‑guesses; decisions can't chain.

Harness — What the agent can read, modify, and how it gets feedback. Without it: Permissions too wide, or agent only outputs text without evidence.

Verification Evidence — How to know this action produced a valid result. Without it: Tests don't run; results only proven by self‑generated summaries.

Human Checkpoint — Which forks involve product, architecture, and responsibility. Without it: Model casually makes irreversible decisions for humans.

Work Site — How the next agent or human picks up. Without it: Long task interrupted, can only archaeologize from chat history.

When You Should Still Plan First

Action loops don't mean everything should start immediately. The author looks at error cost, not task size.

If a change has any of these characteristics, a plan gate remains valuable:

Changes public APIs or cross‑team contracts;

Migrates data with non‑trivial rollback;

Alters permissions, billing, compliance, or security boundaries;

Affects multiple services, clients, or teams;

Each trial is expensive or side effects are irreversible;

Critical verification requires human operation the agent can't do alone.

These tasks need confirmed scope, compatibility, rollback paths, and stop conditions before taking real responsibility; long explanations alone don't solve that.

Here the plan carries governance value. It's the surface for cost confirmation, responsibility boundaries, and high‑impact action approval.

Conversely, these tasks usually fit the action loop better:

Local bug fixes;

Page tweaks within existing patterns;

Adding tests, logs, or copy changes;

Small‑scope refactors that are quickly verifiable and easily rolled back;

Building a prototype first to judge whether a direction is worth pursuing.

Forcing a long plan on such work often puts exploration cost before results. Letting the agent read code, make a small change, run tests, then decide next steps based on evidence is closer to real development.

Therefore the author doesn't favor making Plan Mode a global "press before every task" mode. Safer approach: let it become an on‑demand gate at high‑impact forks.

From a product design view, the agent should judge based on impact scope: which questions can be explored first, which actions need human confirmation once taken; meanwhile keep a plan gate always available so an executing task can switch to human decision at a critical fork, rather than only choosing a mode at the start.

Whether this workflow is truly better can't be judged by how many plan pages it saves. Pick a set of real tasks of similar difficulty, compare full spec vs. short decision record approaches, and measure time from requirement to passing tests, reviewer time to understand key decisions, rework count, and post‑merge issue rate. If the short loop is faster but humans can't explain why a public API changed, it hasn't solved the understanding problem; if a long plan significantly reduces rework on high‑risk migrations, there's no need to cancel it just for the sake of lightweight form.

After Parallel Agents, What Should Humans Watch?

When a single agent modifies a small module, chat history barely serves as process record.

When multiple agents modify the system in parallel, humans can't read every thread and chase every code change reason.

At that point, making plan texts longer adds little value. Humans need an evidence index.

Intent — What is being changed, why it's worth doing

Decision — Which boundary, interface, or assumption changed

Evidence — Which diff, test, log, or screenshot supports the judgment

Status — Verified, pending verification, blocked, or needs decision

Impact — Which modules, downstream dependencies, and rollback paths are touched

The variety of workflows on Hacker News exists because people are looking for different carriers. Some prefer writing plans into the repo, some use task systems to save cross‑session state, some chain multiple short sessions via handoff documents. They aren't necessarily arguing "does planning have value"; they're searching for something more reliable than re‑reading chat logs.

For architects, the review object also changes. Before it was reviewing a proposal, then reviewing code; going forward it's more about what facts this change is based on, which assumptions were overturned by evidence, why this boundary was touched, and which results remain unverified.

Humans don't need to read every model reasoning step, but review must still be able to trace key decisions.

What Plan Mode Should Look Like Next

The author now prefers to see Plan Mode as a combination of three capabilities, not a single UI mode.

First is Exploration Limits . In phases where direction needs confirmation, agents can read code, query docs, run read‑only checks, but cannot produce high‑impact side effects. The key is permissions and tool boundaries, not repeating "don't write code" in prompts.

Second is Decision Records . Agents don't have to generate a long document every time, but must leave current goals, key assumptions, open questions, and reasoning for choices. Records can be short but must not just say "done."

Third is Verification Gates . Whether a task is complete ultimately rests on clear evidence. If tests haven't run, external receipts aren't confirmed, key behaviors lack screenshots or human verification, a model summary alone cannot pass acceptance.

Together, these three give Plan Mode a chance to become a useful work node.

Understand current system
  ↓
Propose current hypothesis
  ↓
Do a low‑risk action
  ↓
Collect external evidence
  ↓
Update goals, constraints, unknowns
  ↓
Pause and request decision at high‑impact forks

This absorbs new facts better than "plan‑approve‑execute" and adds a layer of explainability over "just let the agent run."

Closing Thoughts

The value of the "Plan Mode is dead" title is that it forces us to re‑examine planning's place in agent workflows.

If planning is just giving the model a longer instruction, its value drops as model capability rises.

If planning is about letting humans confirm goals, boundaries, risks, and completion evidence, it still has value — and becomes more important as agents get faster and parallel tasks increase.

The author summarizes the shift in one sentence:

Planning is no longer a text that precedes execution and expires afterward; planning is a control state that updates alongside action.

Low‑risk tasks , with clear tool permissions and verification conditions, can usually let agents advance in the action loop on their own. High‑risk tasks still need human gates at critical forks. Whichever way, the deliverable must not be just a "looks done" summary; it should include goals, decisions, evidence, and remaining risks.

As AI coding accelerates, humans deserve to focus attention on critical decisions: why it did this, what it was based on, what it affects, and when it must stop and wait for a human call.

References

Ayman Nadeem, Plan mode is dead (https://www.aymannadeem.com/artificial/intelligence,/developer/tools/2026/09/24/plan-mode-is-dead.html)

Hacker News, Plan mode is dead discussion (https://news.ycombinator.com/item?id=49840054)

OpenAI, How OpenAI uses Codex (https://cdn.openai.com/pdf/6a2631dc-783e-479b-b1a4-af0cfbd38630/how-openai-uses-codex.pdf)

Anthropic, Claude Code permissions (including plan mode) (https://code.claude.com/docs/en/permissions)

Google, Gemini Code Assist Agent Mode (https://cloud.google.com/gemini/docs/codeassist/agent-mode)

OpenAI, Follow a goal (https://developers.openai.com/codex/use-cases/follow-goals)

Martin Fowler, Harness engineering for coding agent users (https://martinfowler.com/articles/exploring-gen-ai/harness-engineering.html)

Anthropic, Effective harnesses for long‑running agents (https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents)

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI Agentssoftware developmentCodexHuman-in-the-loopClaude CodePlan ModeGemini Code Assistwork loop
Architect
Written by

Architect

Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.