Hierarchical Multi-Agent Architecture: Delegating Control When One Supervisor Isn't Enough
This article explains how hierarchical multi-agent architecture solves the single-supervisor bottleneck by delegating control to intermediate layers that manage local closed loops, using structured task packages downward and decision packages upward, with concrete examples from model evaluation and analysis of trade-offs like constraint dilution and retry multiplication.
Not More Agents, But Concentrated Decisions
A single-layer Supervisor works only while the central agent can understand every result and decide next steps. As the task widens, the center must handle longer contexts, more members, more tools, and validate results against different domain rules. The problem is not token limits but undivided decision boundaries .
Karpathy's context engineering discussion applies: industrial LLM applications must decide what goes into the context window. Adding a sub-agent does not automatically lighten the center; only isolating irrelevant context and professional judgment does. OpenAI's Codex subagents separate instructions, model settings, and tool contexts to keep the main context clean. However, having sub-agents is not the same as having layers . If sub-agents only execute and the top layer still validates each result, it remains a single-layer Supervisor. The intermediate layer must be a control layer that can dispatch, validate, and escalate with evidence when boundaries are crossed.
Hierarchical splits control authority, not headcount. If the middle layer only forwards messages, it adds latency without governance.
Diagram: After the task widens, the center must simultaneously handle context, tool selection, and validation criteria from different domains.
Splitting One Task into Three Local Closed Loops
Returning to the customer-service model evaluation example, the top layer no longer manages every expert but manages three group leads:
Release Owner
├── Performance Lead
│ ├── Benchmark Agent
│ └── Cost Analysis Agent
├── Security Lead
│ ├── Data Flow Agent
│ └── Permission Review Agent
└── Business Lead
├── Customer Service Scenario Agent
└── Fallback Plan AgentThe tree itself has no value; the key is whether each lead can decide next steps within its boundary and own the group's results.
Diagram: Single-layer structure validates each item centrally; layered structure places local dispatch and validation inside professional groups, with the top layer only arbitrating cross-group conflicts.
For example, the performance group receives a load-test record: peak samples show some customer-service sessions exceed latency thresholds, causing the gateway to route requests to fallback-v2. The performance group does not immediately dump all logs to the top; it first runs another sample round to confirm whether the cause is slow model inference, slow retrieval, or premature fallback triggering.
When the performance group reports upward, a single line "latency exceeded" is insufficient. It must specify: which evaluation plan version, which peak sample set, which routing rule enters the fallback, and where the raw report resides.
The security group then examines fallback-v2 and finds it sends session summaries, ticket IDs, and partial user fields to an external service. Returning only "security review passed" leaves the top layer unable to judge. A blocked signal is more valuable: it must carry the triggering plan version, route, fields, and evidence location.
The release owner need not reread both groups' full dialogues but must see the cross-group conflict: the performance fix changed traffic routing, and the new routing triggered a security block. The top layer delegates to the business group to assess whether a local fallback can be used, then asks performance and security to re-verify.
Diagram: Groups handle local issues first; once conditions cross professional boundaries, they escalate with evidence, and after resolution, return to relevant groups for re-verification.
The change after layering: group-internal details are absorbed by the lead; only cross-group conflicts reach the top. The top's scope shrinks, but its responsibility remains.
Task Packages Down, Decision Packages Up
Hierarchical diagrams are usually trees. The real difficulty is not drawing the tree but defining what flows between layers.
Anthropic's post-mortem of their multi-agent research system mentions a game of telephone: information degrades with each handoff. They had some sub-agents write full artifacts to files or an artifact store and pass references back to the coordinator. This lets the upper layer read summaries first and drill into original artifacts when needed.
The ROMA recursive meta-agent paper tackles the same problem: tasks decompose downward into a subtask tree, results aggregate upward with structured compression and validation of intermediate results. Its performance numbers are valid only within its experimental scope and cannot be extrapolated to "all layered systems are stronger." More importantly, it exposes an engineering pitfall: designing decomposition without designing aggregation hides problems inside layers.
Cross-layer information splits into two types: downward task packages , upward decision packages .
A task package must specify goal, hard constraints, input version, budget, permissions, and stop conditions. For example, a task to the security group is not "check data security" but "check whether fallback-v2 sends customer-service session summaries, ticket IDs, or user fields to external services; do not modify the current plan version; return blocked immediately on external egress."
A decision package returns result, evidence references, quality status, unresolved issues, termination reason, and escalation target. It is not a polished "done" note but the basis for the upper layer to continue deciding.
Implementation must preserve at least these fields:
constraints – which hard constraints must not be lost during decomposition and summarization
input_refs – which version of task and evidence this layer actually read
evidence_refs – where the upper layer can trace back data, logs, or documents
quality_status – distinguishing completed, failed, blocked, and pending-confirmation
escalation_target – where the issue goes when this layer lacks authority
Diagram: Task packages constrain scope downward; decision packages carry results and evidence upward; the middle layer must complete a local closed loop first.
The format can be JSON or referenced documents. Format is secondary; the key is summary and evidence must be separated : summaries let the upper layer grasp quickly, evidence lets it retrieve original facts when needed.
If you only decompose tasks downward without designing upward aggregation, more layers produce a chain of ever-shorter summaries.
Every layer reports, but during debugging you cannot find which version a fact came from, which agent produced it, or under what constraints it was validated. This is invisible in demos but becomes critical during troubleshooting.
Layers Distribute Complexity but Amplify Errors
The benefit of layering is that each control layer only needs to understand a smaller scope. But each added boundary adds a chance for constraints to be compressed and results to be misread.
Hard constraints soften during restatement. The top writes "no sensitive fields sent overseas"; the middle simplifies to "check data security"; the bottom validates the wrong scope.
Local passes are mistaken for global passes. Performance and security each pass individually, but that does not mean the combined scenario "peak traffic triggers fallback" passes. Local closed loops resolve group issues; cross-group conditions must return to the upper layer.
Retries multiply across layers. Top retries once, lead retries twice, bottom tool retries three times – actual calls are multiplicative, not additive. Safer: allocate budgets top-down so each layer knows its remaining quota, rather than each setting its own seemingly reasonable limit.
Layering saves the scope a single control layer must simultaneously understand; it adds cross-layer coordination, state alignment, and fault localization. Both sides must be accounted for.
Before Layering, Verify the Middle Layer Can Close the Loop
Deciding whether to use Hierarchical does not start by counting agents. First ask whether the middle layer can do three things:
Independently dispatch within a clear boundary and maintain its local state.
Validate against its professional rules, not just forward raw output upward.
When lacking authority, return with evidence, constraints, and an escalation target.
Only if all three are feasible is the middle layer a control layer; otherwise it is just an extra hop.
If the path is fixed, a Workflow is easier to replay and validate. If the team is small and the top can understand all results, a flat Supervisor is simpler. If the task requires peer agents to continuously hand off control, the problem moves toward the next article's Swarm/Handoff pattern.
Frameworks name "layering" differently. CrewAI calls a Manager overseeing multiple members a Hierarchical Process. Google ADK documents top, middle, and bottom hierarchical task decomposition. LangChain's current migration guide is more direct: flatten when lower layers are independent; keep nested layers only when middle coordination is truly needed.
As of October 7, 2026, the dedicated langgraph-supervisor library is no longer actively maintained. LangChain now recommends most scenarios use a "primary agent invoking sub-agents as tools"; if middle coordination is genuinely needed, wrap the middle agent as a top-level tool. This does not mean layered collaboration is obsolete; it reminds us: layering is a control-authority design, not the name of a particular constructor.
If the system still needs static subgraph discovery, layered checkpoints, or cross-layer shared state, hiding all middle layers inside tool calls makes later debugging painful. Explicit graphs or custom state machines are heavier but keep boundaries clearer.
Hierarchical does not mean more layers equal more professionalism. A middle layer without independent decision space adds waiting and information loss, not governance.
Layering reduces not the total agent workload, but the scope each control layer must simultaneously understand. The price paid is that every cross-layer hop may compress facts, multiply retries, and hide global problems behind local passes.
One step further: if even stable superior-subordinate relationships are removed, letting the current agent hand control to the next, the system moves from layered collaboration to Swarm/Handoff.
References
Andrej Karpathy, X post on context engineering (https://x.com/karpathy/status/1937902205765607626)
OpenAI Developers, Codex subagents discussion (https://x.com/OpenAIDevs/status/2033637455136731431)
Anthropic, How we built our multi-agent research system (https://www.anthropic.com/engineering/multi-agent-research-system)
LangGraph Supervisor, GitHub repo and README (https://github.com/langchain-ai/langgraph-supervisor-py)
LangChain, Subagents (https://docs.langchain.com/oss/python/langchain/multi-agent/subagents)
LangChain, Migrate from langgraph-supervisor (https://docs.langchain.com/oss/python/migrate/langgraph-supervisor)
Google ADK, Multi-agent workflow patterns (https://adk.dev/workflows/patterns/)
CrewAI, Hierarchical Process (https://docs.crewai.com/v1.15.23/en/learn/hierarchical-process)
ROMA: Recursive Open Meta-Agent Framework for Long-Horizon Multi-Agent Systems (https://arxiv.org/abs/2602.01848)
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architect
Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
