How We Cut Token Costs 50% in a Multi-Agent Harness Workflow

A team reduced token costs 50–65% in a 1-TL + 6-sub-agent Harness workflow by applying three principles—load only needed context, eliminate irrelevant context, remove duplicate context—through ten techniques including progressive disclosure, code graphs, CLI over MCP, and parallel tool calls, measured via AgentLens.

ITPUB
ITPUB
ITPUB
How We Cut Token Costs 50% in a Multi-Agent Harness Workflow

Background and Challenge

The team used AI agents to drive full-stack development—from requirements analysis to automated testing—orchestrated by a tech-leader harness skill comprising 1 TL (Task Leader) and 6 sub-agents:

TL (orchestration)
├── Wave 1 (parallel): Backend Agent + Frontend Agent
├── Wave 2: Quality Review Agent
├── Wave 3: Test Agent (case generation + execution)
├── Wave 4: Visual Validation Agent
└── Wave 5: Agent Evaluation

A medium-sized requirement typically runs 5–6 waves, 20+ sub-agent invocations, and hundreds of tool calls, making token cost a critical issue.

Pain Point: Six Cost Sources

Token consumption per task was categorized into six sources (system prompts, tool returns, history messages, etc.). Without per-wave cost breakdown, the team could not identify where money was being spent.

Measurement First: AgentLens

Using CodeBuddy (internal IDE) which reports per-turn tool calls and token usage to the AgentLens platform, they traced single calls by TraceId and aggregated full requirements by SessionId. This visibility revealed the cost distribution and set the optimization target.

Three Guiding Principles

Let AI see only the context needed right now — load on demand.

Reduce irrelevant context — keep unrelated content out from the start.

Reduce duplicate context — avoid re-charging for the same content across turns.

Ten Optimization Directions

3.1 Let AI See Only Current Needed Context

3.1.1 Progressive Disclosure

Adopted Anthropic's progressive-disclosure architecture: three layers (core skeleton, conditional resources, step details). Only the skeleton loads initially (~1,000–2,000 tokens even with 20 skills), cutting context ~90% vs. monolithic prompts. Moved conditional content (e.g., 40-line spec template, 38-line DB rules) and all step details into references/ resources; the skill file keeps only the skeleton. The automation test skill shrank from 198 to 128 lines (-35%).

3.1.2 Deterministic Operations via CLI Scripts

Replaced ad-hoc command construction with a dev-env.sh script handling DB migration, compile, service start, health checks. AI only passes parameters, eliminating token-heavy trial-and-error.

3.1.3 Prefer CLI over MCP

MCP tool calls incur hidden costs: LLM decision + tool JSON schema (10–15 KB/turn) + large tool results + another LLM turn to process results. Switched Playwright MCP to Playwright CLI: AI generates spec files, CLI executes them in batch. This separates reasoning (AI) from execution (script), saving tokens and time, and enables parallel test runs.

3.1.4 MCP Data Fetching via Sub-Agents

Originally the long-lived TL called TAPD/Figma MCP directly, retaining massive payloads (thousands of lines) for the entire workflow. Created dedicated sub-agents ( tapd-req-analyzer, figma-design-analyzer) that return only structured summaries. Measured single-session input tokens dropped from 1,030,000 to 634,905 (-38.4%); real savings compound over subsequent turns.

3.1.5 Long-Term Memory On-Demand Indexing

The context-keeper skill previously loaded entire matched documents. Introduced an INDEX.md (titles, tags, summaries). Flow: read index → keyword match + category weighting → read top-3 full documents only. Index is tens of lines vs. dozens of full documents, avoiding "loaded but unused" tokens.

3.2 Reduce Irrelevant Context

3.2.1 Single Agent → Multi-Agent Split

Original monolithic agent mixed frontend/backend/test/visual concerns in one ever-growing history. Split into TL + role-specific sub-agents ( backend-dev, frontend-dev, test-runner, visual-reviewer, code-reviewer). Each sub-agent carries only its domain prompt/tools, dies after its wave, and Wave 1 runs in parallel. Splitting itself adds overhead (6 system prompts), so they first predict requirement size (S/M/L); only M/L trigger multi-agent mode.

3.2.2 Agent-Specific Configuration

Defined custom agents with frontmatter tool allowlists (e.g., backend-dev gets only file/exec tools, no MCP servers). This removes unused tool schemas from context. Also moved stable dispatch instructions into sub-agent system prompts (better cache hit), shrinking TL dispatch prompts from 10–15 lines to 2 dynamic lines. Enabled model routing: test/visual agents use GLM-5v (~36% of Sonnet cost), yielding -64% cost for those roles.

3.2.3 Code Graph Replaces Blind Search

Integrated graphify (AST+semantic index). Instead of iterative search_content returning many irrelevant matches, agents query the graph to pinpoint files, then read_file only those. Official estimate: tens of times token reduction. Empirical test on same task: total tokens 875k → 677k (-22.7%), input tokens -22.8%. Fewer exploration rounds also cut API calls.

3.3 Reduce Duplicate Context

3.3.1 Stable Prefix Design (Prompt Cache)

LLM KV-cache reuse (≈10% price) requires identical prompt prefixes. Fixed two prefix breakers: (1) interleaved static/dynamic content in dispatch prompts → moved all dynamic content to the end; (2) TL accumulated full progress dashboards in history → externalized progress to a file (read on wake, single-line stage updates). Also enables session resume.

3.3.2 Avoid Reloading Skills

Frontend agent was re-invoking context-keeper skill (200 lines) despite TL already having loaded experience and written the tech spec. Removed the redundant call; downstream agents read the spec document instead.

3.3.3 Compress CLI Output with rtk

Integrated rtk (CLI output compressor) via CodeBuddy PreToolUse hook. Official claims 60–90% reduction; actual varies by command (e.g., ps aux -98.9%, git status -31%). Configured globally in ~/.codebuddy/settings.json. Evaluation via rtk gain or direct CLI diff (deterministic) rather than full workflow re-runs (non-deterministic noise).

3.3.4 Parallelize Independent Tool Calls

Identified silent serializations: TAPD + Figma summaries (no data dependency) → now dispatched in same message via two Task calls. Test execution: Playwright CLI runs multiple specs with internal workers; LLM side makes 1 call + reads summary. Rule: if no data dependency, parallelize — saves history re-packaging rounds.

Results

Per-wave measured reductions aggregated to an estimated 50–65% full-flow token cost reduction for medium requirements.

Key Takeaways & Rollout Priority

Token saving ≠ feature loss — only changes when and how context loads.

Collect once upstream, pass via documents — biggest waste is every agent rediscovering the same info.

Cheapest call is no call — use CLI/pre-fetch for deterministic work, reserve LLM for semantic tasks.

Parallelize whenever possible — saves not just latency but repeated history packaging.

Rollout order: size prediction + agent split (architectural prerequisite) → measurement + skill reorder + global rtk (quick win) → conditional content extraction, state externalization, parallelization audit → code graph, CLI-over-MCP, tool allowlists, sub-agent MCP, memory indexing (systematic mid/long-term).

References

Anthropic Multi-Agent Research System: https://www.anthropic.com/engineering/multi-agent-research-system

GitHub Copilot Token Efficiency: https://github.blog/ai-and-ml/github-copilot/improving-token-efficiency-in-github-agentic-workflows/

rtk CLI Compressor: https://github.com/rtk-ai/rtk

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

multi-agentprogressive disclosuretoken optimizationLLM cost reductioncode graphHarness workflowAgentLensMCP vs CLI
ITPUB
Written by

ITPUB

Official ITPUB account sharing technical insights, community news, and exciting events.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.