AWS Strands Harness Cuts Agent Costs 77% by Swapping Runtime, Not Model
AWS open-sourced Strands Harness, a pre-assembled agent runtime that reduces costs 77% versus Claude Code on the same Claude Fable 5 model by optimizing context management, memory layering, prompt caching, and progressive skill loading, shifting evaluation focus to Model × Harness × Workload combinations.
AWS released Strands Harness on September 21, 2026, accompanied by a self-run benchmark comparing six agent harnesses — Strands Harness, Claude Code, Codex, OpenCode, oh-my-pi, and DeepSeek Harness — all using the identical underlying model, Claude Fable 5, across six benchmarks including Terminal-Bench 2.1, ALFWorld, ContextBench, GAIA, WebShop, and τ³-bench. The results show Strands Harness completed 89 trials at $56.29 with an accuracy of 69.7, while Claude Code cost $248.05 for 61.8 accuracy — a 77% cost reduction and 7.9-point accuracy gain. DeepSeek Harness was cheaper at $40.30 but scored only 59.5, illustrating that token efficiency alone does not guarantee task success.
Context Management: 1500-Token Truncation and 85% Compaction
Strands achieves its token efficiency through two default ContextManager rules. First, any tool result exceeding approximately 1500 tokens is truncated: only a head/tail preview remains in the active context, while the full output is written to an external stash and can be retrieved via a retrieve_context tool call. This converts what would otherwise be repeated input tokens across multiple model turns into a one-time write plus on-demand reads. Second, when context window utilization reaches 85%, a summarization compaction runs, preserving the most recent messages (e.g., the last 4) while compressing older history. If compaction still leaves the context overfull, a Context Recovery mechanism continues the agent loop instead of aborting with a context-length error. The article illustrates the compounding effect: a tool returning 4000 tokens over 20 rounds would accumulate 80,000 tokens of tool output if retained naively, forcing the model to re-read large portions each turn.
Layered Storage: Context Window, Session, and Long-Term Memory
Strands separates state into three tiers. The context window acts as working memory, aggressively managed via offloading and compaction. Session state persists the full conversation to disk under ./.agent/sessions keyed by session ID, enabling task resumption after process restart. Long-term memory is handled by a MemoryManager with three capabilities: Recall (on-demand search of historical knowledge), Injection (prepending relevant memories to the prompt before a run), and Extraction (deriving durable facts from conversations and writing them to a memory store). Recall and Injection are enabled by default when a memory store is configured; Extraction requires explicit opt-in. The article analogizes this to traditional storage hierarchy: context window ≈ RAM, session ≈ task checkpoint, memory store ≈ persistent knowledge base.
Skills and Subagents: Progressive Disclosure
Instead of loading all skill instructions into the system prompt at startup, Strands uses progressive disclosure: only skill names and descriptions enter the initial prompt; full instructions are fetched via tool call when the model decides they are needed. This mirrors the tool-result offloading philosophy — avoid pre-loading potentially irrelevant tokens. The default create_harness() bundles shell, read, write, edit, web_fetch, web_search, programmatic_tool_caller, and subagent tools, plus todos, environment plugins, long-term memory, prompt caching, context manager, and configurable MCP servers. A generalist subagent handles open-ended subtasks while the main agent tracks progress via todos.
Prompt Caching and the True Cost Equation
Strands enables prompt caching in auto mode, leveraging provider-level caching (e.g., Anthropic, Bedrock) for system prompts, tool definitions, and stable message prefixes. The article presents a cost formula: Effective Context × Call Count − Cache Hits − Offloaded History + Compaction Overhead + Subagent/Tool Call Costs . Because caching capabilities differ across providers, the same harness can yield different bills on different backends, reinforcing the need for harness-level benchmarking.
Model × Harness × Workload: The New Evaluation Unit
Strands standardizes a Model Interface that unifies streaming, tool calling, and structured output across providers (Bedrock, Anthropic, OpenAI, Google, Ollama, LiteLLM, Mistral, Llama API, llama.cpp, Vercel). This decouples the model from the harness, and the harness from the cloud platform (runnable on any Linux container host: Cloudflare Containers, Azure Container Apps, Google Cloud Run, Amazon ECS, Bedrock AgentCore, Modal). However, provider differences in prompt caching, tool schema support, reasoning effort, context window, and parallel tool calls mean swapping models does not guarantee identical performance. The AWS benchmark itself demonstrates that Model × Harness is an interacting pair: the harness determines what the model sees and when it acts, while the model’s capability determines how well it uses that information. Consequently, the article argues that production evaluation should target Model × Harness × Workload triples — measuring cost, success rate, and step count on realistic tasks — rather than model leaderboards alone.
Conclusion
Strands Harness represents a shift in agent scaling: from improving the model alone to hardening the runtime that surrounds it. The benchmark data is AWS-internal and not yet independently verified; DeepSeek Harness shows a lower-cost but lower-accuracy alternative. Nonetheless, the direction is clear — models set the capability ceiling, but context, memory, tools, skills, cache, and execution loops determine how much of that ceiling is realized. AWS aims to fix the latter stack so developers can swap "brains" without rebuilding the "body."
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunTalk
Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
