Inside Pi: Minimalist Agent Harness Design & Core Loop Architecture

The article analyzes Pi's agent harness architecture, showing how its core loop stays minimal by handling only model-tool-result cycles while pushing session management, MCP integration, Codemode orchestration, and security boundaries to the edges, enabling extensibility without complicating the main execution path.

Architect
Architect
Architect
Inside Pi: Minimalist Agent Harness Design & Core Loop Architecture

The article examines the Pi agent harness (main branch commit 96abe9a, version 1.1.0) by tracing a single task through its source code to reveal how minimalism is achieved not through small code size but by keeping the core execution path short and decision-free.

Core Loop and Layered Architecture

Pi's codebase is organized in clear layers: pi-ai — abstracts differences between model providers. pi-agent-core — implements the agent loop. pi-coding-agent — adds file tools, session management, context compression, extensions, and the TUI.

The author focuses on pi-agent-core, which knows nothing about Git, Markdown, MCP, terminals, or session files. It only holds the model, messages, and tools. A task follows this minimal path:

User message → Model response → Tool calls → Tool results → Back to model → End
Pi's Harness layering
Pi's Harness layering

Despite its brevity, the core loop handles critical edge cases:

Parallel tool execution: Multiple tool calls in one response run in parallel by default; if any tool declares itself serial, the whole batch runs sequentially. Results are written back in the original call order, preserving the model's intent while avoiding waiting for slow tools.

Truncated model responses: If a response is cut off mid-JSON, Pi marks the call as failed and returns it to the model for retry, rather than executing malformed parameters.

User interruptions: A steering message (e.g., "don't touch config") is injected after the current tool batch finishes; a follow-up message (e.g., "add tests") starts a new turn after the current one ends naturally.

The core loop hard-codes none of these product policies. Instead it exposes three well-defined extension points: prepareRequest() — shapes input before calling the model. prepareNextTurn() — decides how to continue. finishTurn() — judges whether to stop or continue.

Context compression, model switching, and extension logic plug into these hooks, while the main loop still only manages "model — tool — result".

Model-tool loop in one task
Model-tool loop in one task

Tools as Default Vocabulary, Not Capability Ceiling

pi-coding-agent

currently defines eight built-in tools: read, bash, powershell, edit, write, grep, find, ls.

The default coding set is read, bash, edit, write; a read-only set ( read, grep, find, ls) is also provided. Users and extensions can replace these defaults.

A crucial distinction: the Pi process's capabilities and the tools visible to the model in a given turn are not the same table. When the tool set changes, declareToolChanges() records additions and removals as session events, enabling full restoration on session resume. Tools are part of task state, not static system-prompt text.

MCP extends this idea. A server may expose dozens of tools, but they need not all appear in the model's schema. Pi offers four exposure modes for MCP tools: direct — declared directly to the model (high-frequency tools). deferred — searched and loaded on demand. codemode (default) — invoked via scripts inside Codemode. hidden — available to the runtime but not exposed to the model.

With codemode as default, the model first discovers capabilities via searchTools and describeTool, then calls them from a script. This keeps the context small. For ordinary MCP calls, text results exceeding 20 KB are truncated (head/tail kept, full output saved to a temp file). Codemode scripts receive the full payload, can filter/aggregate, and return only the essential summary.

System prompts follow a similar sectional composition: tools, rules, project context, skills, and working directory each form independent sections. Changes are recorded as patches in the session; the runtime assembles the complete prompt for each request, but the composition and its history remain traceable and restorable.

Capability access vs model visibility
Capability access vs model visibility

Session Persistence and Projection

Long tasks face a tension: the session log should be complete, but the model context cannot grow indefinitely. Pi separates these concerns.

Storage: Append-only JSONL tree. Every record has an id and parentId. Continuing from an old node creates a new branch; switching branches merely moves the leaf pointer. Original records are never deleted or mutated.

Projection: buildSessionProjection() generates the model's view from the current branch. It walks back to the latest compaction, then reassembles the summary, retained messages, and compressed new messages into a single context. Early raw records stay on disk but are excluded from the request.

Corrections: To alter a past message's effect on context, Pi appends a context edit describing the replacement/removal, leaving history intact.

Compaction summaries also store lists of files read and modified, giving the agent a persistent memory of touched artifacts even after raw tool outputs are dropped.

Before each model call, AgentSession swaps its in-memory message list with the projection. On context overflow, Pi auto-compacts and retries once; a second failure ends the task cleanly.

Session preserves facts; projection serves the present. Each has a clear goal and they do not compromise each other.

Codemode: Orchestrating Multiple Tool Calls in One Turn

As tool counts grow, lazy schema loading is not enough. The bigger cost is intermediate results polluting the conversation. Example: the model must query dozens of issues, filter by label, recency, and assignee. Calling tools one by one forces the model to re-read large intermediate payloads just to produce a short list.

Codemode lets the model emit a JavaScript snippet that runs in a QuickJS VM. The script can invoke multiple tools, filter, and aggregate, returning only the final list to the main context. Nested tool calls do not become new conversation messages; the runtime logs only bounded metadata (name, arguments, status, duration, error).

The VM is deliberately restricted: no Node APIs, no fetch, process, require, module system, or WebAssembly. Data crosses the VM boundary via JSON serialization. Execution has time limits, memory limits, and an interrupt mechanism for runaway scripts.

However, QuickJS solves orchestration and context pollution — it does not create a security sandbox. When a script calls read, bash, write, or an MCP tool, the actual execution happens in the host process with whatever permissions the Pi process holds. The VM's lack of process does not strip permissions from the host-side bash tool. Codemode primarily shortens the data path between model and tools; security boundaries must still be enforced at the tool or runtime level.

What Pi Leaves at the Boundaries

No per-call confirmation and no built-in permission system (directory, command, or network allow-lists). Pi inherits the launching user's and host process's permissions.

Project trust controls whether an unfamiliar repository can silently load its own extensions, skills, and configs. It does not restrict what an already-enabled bash or write tool can access. AGENTS.md and CLAUDE.md are loaded by default unless context loading is explicitly disabled; trust is a loading policy, not an OS security boundary.

Extensions run in-process with full host access. For untrusted code, secrets, or multi-tenant workloads, the recommended approach is full process isolation or routing high-risk tools to containers/remote sandboxes (Plain Docker, Docker Sandboxes, OpenShell, Gondolin). The choice and granularity are left to the operator.

No built-in sub-agent or plan-mode workflows. Teams needing approvals, unified policies, collaboration, or centralized auditing must build those on top of Pi.

The repository now includes pi-durable and a set of server/protocol/client packages for persistent runtimes, but the standard Coding Agent's main path still revolves around pi-agent-core and AgentSession. Durable integration lives in experimental. Complex capabilities have their own space without being forced into the default CLI.

Conclusion: Minimalism Is Fewer Decisions on the Core Path

Re-reading Pi, the author concludes that a harness's minimalism cannot be judged by repo size, system-prompt length, or default tool count. The better measure is how many decisions the core path makes for upper layers.

Agent Core guards only the model-tool-result loop.

Session stores everything; projection serves the current branch.

Tools can be numerous, but the model sees only what it needs right now.

Complex orchestration moves into Codemode.

Security isolation stays at the tool and runtime level.

Each layer is transparent; together they support real long-running tasks. This trade-off suits teams that want a clear foundation, not a fully governed platform. If fine-grained permissions, mandatory approvals, and centralized governance are required, Pi provides only the groundwork. As a harness, it cleanly separates "how it runs by default" from "how to extend when needed."

The author now evaluates agent frameworks by first finding their most-traveled path and then checking whether that path has been dragged into an all-knowing controller as features accumulate. Pi's repository has grown, yet its main loop remains readable end-to-end — that, to the author, is its most valuable minimalism.

References

Pi official repository: https://github.com/earendil-works/pi Pi documentation (How Pi Works, MCP Servers, Codemode, Containerization): https://pi.dev/docs/latest Mario Zechner, "What I learned building an opinionated and minimal coding agent": https://mariozechner.at/posts/2025-11-30-pi-coding-agent/ Armin Ronacher, "Pi: The Minimal Agent Within OpenClaw": https://lucumr.pocoo.org/2026/1/31/pi/ Armin Ronacher, "What is Codemode":

https://lucumr.pocoo.org/2026/10/6/codemode/
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI AgentsMCPSession ManagementPiMinimalist DesignAgent HarnessCore LoopCodemode
Architect
Written by

Architect

Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.