SwarmAgent: Turning Multi-Agent Concepts into a Verifiable MVP

SwarmAgent evolves from a conceptual multi-agent template into a verifiable MVP by implementing true parallel execution, reliable result passing, standard MCP compliance, security boundaries, and comprehensive contract tests, while honestly documenting its current limitations.

Chengwu Tech Stack
Chengwu Tech Stack
Chengwu Tech Stack
SwarmAgent: Turning Multi-Agent Concepts into a Verifiable MVP

From Concept to Verifiable MVP

The author revisits the open-source project SwarmAgent, which aims to use a Supervisor for task decomposition and aggregation, with specialized agents (Researcher, Writer, Coder) collaborating via Web UI, SSE streaming, file sessions, a shared blackboard, and MCP tools. While the project appeared feature-complete, the author emphasizes that "looks runnable" differs from "verifiably works" by a full engineering contract.

Three Common Multi-Agent Shortfalls

1. Fake Parallelism

Many projects claim parallel execution via spawn_agent but actually run tasks sequentially. If two independent tasks each take 10 seconds, sequential execution still needs 20 seconds. SwarmAgent now defines true parallel semantics: when the model issues multiple spawn_agent calls in the same turn, they run concurrently in the same event loop and return results in original call order after all complete. The team verified this with real timing tests confirming temporal overlap, not just async keywords.

2. Results Not Actually Flowing Between Agents

Logs may show Researcher, Writer, and Supervisor taking turns, but the Writer might not receive the Researcher's findings, and the Supervisor might hallucinate a summary. SwarmAgent uses two paths for artifact transfer:

Tool call results return directly to the Supervisor's context.

Structured artifacts are written to a session-level shared blackboard and injected into downstream agents.

An end-to-end test with a deterministic model stub validates the full chain:

Researcher researches
↓
Writes to shared blackboard
↓
Writer reads research and drafts
↓
Supervisor aggregates final answer

The test explicitly asserts that the Writer's input contains the Researcher's actual output.

3. MCP Configuration ≠ MCP Compliance

Listing MCP servers in config doesn't prove runtime protocol adherence. SwarmAgent replaced its hand-rolled subprocess + JSON-RPC approach with the official MCP Python SDK, executing the full lifecycle:

Start stdio server
↓
initialize
↓
list_tools
↓
Confirm tools exist
↓
call_tool
↓
Close session and subprocess

This yields clearer error boundaries: connection failures, missing tools, and server errors are caught at the correct protocol layer. Default wrappers cover file reading, web fetching, and Brave search; third-party MCP servers still require a live smoke test before deployment.

Security Boundaries for Network-Exposed Agents

Once bound to a public address, file sessions, context injection, and control interfaces become real attack surfaces. Four basic boundaries were added:

Optional API Key: Setting SWARM_API_KEY requires X-API-Key or Bearer token for all /v1/* endpoints; health checks and static homepage remain public. The Web UI stores the key only in the browser's sessionStorage.

UUID Session Boundary: External session IDs must be UUIDs, preventing arbitrary strings from entering file paths. Non-existent session updates return 404 instead of silently creating directories.

Approval Timeout Defaults to Deny: High-risk operations requiring human approval fail closed on timeout; "auto-approve on timeout" would turn the approval gate into decoration.

Markdown Output Sanitization: Model outputs and MCP returns are treated as untrusted. The Web UI passes all Markdown through DOMPurify before DOM insertion, covering historical messages, streaming messages, and final results to prevent XSS.

Why It's Not a "Production-Grade Framework"

The author stresses that an open-source project's value lies in defining boundaries, not listing capabilities. SwarmAgent is suitable for:

Learning Supervisor–Specialist collaboration patterns.

Rapidly validating multi-step workflows (research, writing, review).

Serving as a lightweight prototype for internal agent products.

Extending with custom roles, tools, sessions, and front-end interactions.

It currently lacks:

Distributed task queues.

Task recovery after process exit.

Redis or database session backends.

Full command-execution sandbox.

Multi-tenant permissions and audit.

Long-running MCP connection pools and production monitoring.

It is not a replacement for LangGraph, AutoGen, or OpenAI Agents SDK, nor an out-of-the-box enterprise platform. Its value is compressing a multi-agent execution chain into something small, clear, and verifiable so developers can read, verify, and replace each link.

Proving the Upgrade Isn't Empty

New MVP contract tests cover:

Real parallel timing of multiple spawn_agent calls.

MCP SDK initialization, tool discovery, and invocation.

Approval timeout fail-closed behavior.

API Key protection of status endpoints.

UUID session file boundaries.

Markdown rendering safety.

Full Researcher → Writer → Supervisor collaboration chain.

Current local verification results:

Pytest: 34 tests all passing.

Original R1–R8 validations: 8 rounds all passing.

Ruff and Python compilation checks pass.

Application code coverage: 66%.

No suspected API keys, tokens, or private keys in plaintext.

GitHub Actions CI now runs on Windows and Linux with Python 3.11 and 3.12. Coverage is still insufficient; third-party MCP and Docker need further real-environment validation — these gaps are explicitly documented in the project docs and roadmap.

Closing Thoughts

Judging an agent project's value shouldn't depend on role count or complex README diagrams. Four things matter: honest scope, readable code, runnable verification, and test-proven collaboration chains. SwarmAgent is small but now honest, readable, runnable, and test-backed. Developers researching multi-agent systems or building lightweight prototypes without heavy frameworks can start here — and are encouraged to validate it with real business scenarios rather than adding more impressive-sounding agent names.

Project repositories:

https://gitee.com/hmk_855_admin/swarmagent

https://github.com/hanmengkai/swarmagent

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

MCPopen-sourceMVPmulti-agentcontract testingparallel executionsecurity boundariesSwarmAgent
Chengwu Tech Stack
Written by

Chengwu Tech Stack

A powerful mindset is a lifelong treasure!

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.