Pi vs DeepSeek Harness: Agent Loop Design for Continuation, Closure & Recovery

This article compares Pi and DeepSeek Harness agent loop architectures across four control points—input attribution, tool scheduling, closure permissions, and persistent evidence—revealing how each handles mid-task user constraints, parallel tool execution, token truncation, and crash recovery through distinct runtime contracts and durable extensions.

Architect
Architect
Architect
Pi vs DeepSeek Harness: Agent Loop Design for Continuation, Closure & Recovery

Two Control Planes at Different Layers

Both Pi and DeepSeek Harness (DSH) implement an agent loop, but the control plane sits at different layers. Pi's base loop keeps the short loop and extension points in the core, delegating input arrangement, lifecycle, and callbacks to the caller. Pi Durable adds a persistent runtime layer (session, task checkpoint, inbox, replay) on top without changing the base loop. DSH embeds inbox projection, event boundaries, and scheduling results directly into the loop runtime, recording turn/start, step/start, step/end, and turn/end events with termination reasons.

Four Hierarchical Levels

The article distinguishes four levels that must not be conflated:

request – a single provider call from send to stream end.

step – one model response plus the tool calls and results it initiates.

turn – a group of steps around a set of inputs; decides whether the model must return again.

activity – whether the agent still has running work or has returned to an idle state ready for new tasks.

A step may finish while the turn still has a next-step message; a turn may end while tool processes, event listeners, or cleanup tasks are still running. Pi's agent_end, DSH's turn/end and whenIdle() observe different levels and carry different guarantees.

Pi: Small Core Loop, Extension Points at Boundaries

Input Attribution: steering and follow-up Division

steering

belongs to the current work. If a user says "don't modify config files" while the model processes files, the message enters the next model request after the current tool batch completes. Already-started tools are not revoked; the model's next visit sees the new constraint. follow-up belongs to the next turn; it queues after the current turn ends naturally. Two queues separate "correction for current request" from "next task", so the caller need not guess message ownership from callback order.

Tool Batches: One Result Cannot Close the Whole Batch

Pi only treats a tool batch as eligible for early termination when all completed results in that batch carry terminate: true. A single parallel tool result cannot decide the batch behavior. A search tool may return data; only a transaction-committing tool may express "no need to ask the model again this turn." The newer finishTurn hook runs after assistant messages and tool results are processed, before emitting turn_end. It can return { action: "end" } to end the turn, { action: "continue" } to guarantee at least one more model request, avoids duplicate requests when inputs or results already schedule one, and treats error and aborted as hard exits. "No tool call means end" no longer holds; finishTurn decides jointly with existing inputs and tool scheduling.

Truncation and Lifecycle: Prefer Leaving Failed Results

When model output hits the token limit, Pi does not execute tools even if tool calls are visible; instead it synthesizes a failed result for each call so the next request can resend full parameters. Half-parsed JSON does not equal complete intent; refusing execution is safer for file writes, HTTP requests, or database mutations. agent_end does not mean all async work is drained. Subscribers' async handling of agent_end must finish before the agent enters idle. Resource cleanup, orchestration tasks, and whenIdle() must respect this lifecycle boundary.

Pi Durable: Persistent Runtime Above the Base Loop

@earendil-works/pi-durable

provides durable session, task checkpoint, session inbox, tool replay strategies, and process-restart recovery. Describing Pi as "only saves session history, recovery depends on host" is outdated. The base package handles the short loop and extension points; the durable harness runs a separate runtime layer that persists session, task, and recovery facts.

DSH: Every Continuation Leaves Rebuildable State

In DSH, every turn, step, and inbox change is accompanied by events. A turn begins with turn/start, each model call wraps with step/start and step/end, and the turn ends with turn/end recording the termination reason. Inputs are not just in-memory arrays; session events project an inbox.

Inbox: Separating Current Step from Next Turn

The inbox has two entrances: next-step and next-turn. When claiming a step, the system consumes all next-step messages; if this is a new turn, it also takes one next-turn. Tool context, plugin messages, and user next-turn input never mix due to callback ordering. Insert, replace, delete, and cancel operations are recorded as agent/inbox/spliced events. After a process restart, the system re-projects pending messages from these events. What is saved is not "whether a function was called" but the input state the model's next request must see.

Closure: Tools Propose, Queue Retains Veto

The rules split into two debts:

Model debt : tools have executed; does the model owe another visit?

Message debt : does next-step still hold steering, tool context, or plugin-appended messages?

A tool result's concludesTurn can waive the default model revisit, reducing model debt; if next-step already has messages, the loop still runs another step. Tools can signal "this result is enough to end" but cannot erase already-queued input. agent/turn-stopping fires only when a turn has an end candidate and next-step is empty. If a plugin wants to continue, it writes a consumable message to the inbox; the loop then re-checks the queue. Continuation has events and messages as evidence, not a vague callback return value.

Tool Scheduling: Parallel Execution, Ordered Submission

DSH distinguishes parallel-safe calls and exclusive barriers via tool declarations. The scheduler allows safe overlap but submits results in model-call order, keeping context order stable. On cancellation, started calls are drained; not-yet-started calls receive ABORTED_BEFORE_DISPATCH results. Scheduler failure or step crash before result submission marks unresolved calls with TOOL_OUTCOME_UNKNOWN recovery results. The next request sees paired calls and results, not ambiguous half-records. When output hits max-tokens, DSH marks the turn max-tokens; subsequent steps cannot rewrite it to ordinary completed. The truncation signal persists so external policies can decide retry, manual review, or task termination.

Same Task, Two Loop Executions

At closure, Pi coordinates via steering, follow-up, and finishTurn; DSH coordinates via next-step, next-turn, concludesTurn, and turn-stopping. Pi pushes decision points to the caller; DSH writes the decision process into persistent state.

Point-by-Point Comparison of Four Control Points

Input Attribution : Pi uses steering for current work and follow-up for next turn; Pi Durable persists session inbox. DSH uses next-step and next-turn with event projection.

Tool Closure : Pi examines tool batch results and finishTurn; Pi Durable adds task and replay policy. DSH's concludesTurn only reduces model revisits, cannot skip queued messages.

Concurrency and Cancellation : Pi base package only defines loop boundaries; policies are caller's responsibility. Pi Durable handles via durable tasks and replay strategies. DSH manages exclusive barriers, ordered submission, and post-cancellation result completion.

Recovery Evidence : Pi base makes no full persistence guarantee; Pi Durable saves session, checkpoint, inbox, replay state. DSH saves turn/step events, inbox splices, and unknown-result states.

For short tasks that complete normally, both loops look similar. Differences surface when tools have side effects, processes restart, or tasks queue—falling on recovery evidence and operational cost.

Where Complexity Lives

Pi's base loop advantage: fewer boundaries, faster workflow changes. The caller decides permissions, sandboxing, retries, and external task integration. The trade-off: processes, connections, queues, and side effects created by extensions default to host responsibility; durable recovery requires adopting Pi Durable or another runtime. DSH fixes more control points in the core: inbox has persistent projection, turn/step have event boundaries, scheduler handles concurrency and cancellation, failure results have replayable representations. Long tasks, multi-plugin setups, and recovery audits gain consistent semantics; plugin authors must understand next-step, next-turn, wakeup, event projection, and post-cancellation result states. Pi leaves complexity to the integrator; DSH concentrates complexity into the runtime contract. Pi Durable adds a persistent boundary outside the base loop.

Four Questions for Selection

When designing an agent loop, ask:

Input Attribution : When a message arrives, does it belong to the current step, the current turn's next step, or the next turn?

Model Debt : When must tool results return to the model? Which end signals skip which request? At token truncation, which calls must never execute?

Closure Permission : Do tools only suggest ending, or do they have authority to end? Can queue, policy, or human confirmation veto? For parallel tools, is it any-one-satisfies or all-must-satisfy?

Persistent Evidence : After a crash, can you rebuild the messages the model saw, tool calls and results, cancellation reasons, unfinished tasks? Can irreversible side effects be distinguished among "not started", "execution interrupted", and "result unknown"?

If these four questions lack answers, adding retries, plugins, or planning steps only postpones uncertainty to harder-to-debug positions. Designing an agent loop means turning both "continue" and "end" into state transitions. Pi lets the integrator decide which states need persistence; DSH writes more states into the runtime. Short tasks: evaluate loop malleability. Long tasks: evaluate failure scene reconstruction. The two-layer difference ultimately lands on maintenance cost.

References

DeepSeek Harness official repository: https://github.com/deepseek-ai/deepseek-harness

DeepSeek Harness Agent Loop documentation: https://github.com/deepseek-ai/deepseek-harness/tree/main/packages/core/agent-loop

Pi official repository: https://github.com/earendil-works/pi

Pi Agent Loop source: https://github.com/earendil-works/pi/blob/main/packages/agent/src/agent-loop.ts

Pi Durable Harness documentation: https://github.com/earendil-works/pi/tree/main/packages/durable

Pi containerization documentation: https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/containerization.md

Mario Zechner: What I learned building an opinionated and minimal coding agent: https://mariozechner.at/posts/2025-11-30-pi-coding-agent/

X discussion on Pi vs DeepSeek Harness architecture: https://x.com/mylifcc/status/2087896239455310322

X discussion on DSH vs Pi architectural trade-offs: https://x.com/limboai/status/2087905905664973922

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Event SourcingAgent ArchitecturePiCrash RecoveryAgent LoopTool SchedulingDurable ExecutionDeepSeek Harness
Architect
Written by

Architect

Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.