AWS Strands Harness Slashes Agent Costs 77% Without Changing Models

AWS open-sourced Strands Harness, an agent runtime that cuts costs from $248 to $56 and boosts accuracy on Terminal-Bench 2.1 by optimizing context management, prompt caching, and layered memory, proving agent performance hinges on the harness not just the model.

DataFunTalk
DataFunTalk
DataFunTalk
AWS Strands Harness Slashes Agent Costs 77% Without Changing Models

Benchmark: Same Model, Different Harness, 77% Cost Reduction

AWS ran a harness benchmark across six suites (ALFWorld, ContextBench, GAIA, WebShop, τ³-bench, Terminal-Bench 2.1) using EC2 distributed infrastructure. All systems used the same underlying model: Claude Fable 5. Over 89 trials, Strands Harness cost $56.29 and scored 69.7 accuracy, while Claude Code cost $248.05 and scored 61.8. Oh-my-pi matched Strands' accuracy (69.7) but cost $86.83. DeepSeek Harness was cheapest at $40.30 but scored only 59.5. Across all six benchmarks, Strands averaged ~28% lower token cost than alternatives with equal or higher accuracy.

Table 1: Harness Benchmark Comparison (89 trials, Claude Fable 5) Strands Harness – $56.29 – Accuracy 69.7 Oh-my-pi – $86.83 – Accuracy 69.7 OpenCode – $73.42 – Accuracy 66.3 Claude Code – $248.05 – Accuracy 61.8 DeepSeek Harness – $40.30 – Accuracy 59.5

The data comes from AWS's own testing; the full paper is not yet public, so results should be treated as indicative rather than independent benchmark conclusions. Nevertheless, they highlight that the harness layer can create multi-fold cost differences even when the model is identical.

Context Management: 1500-Token Truncation and 85% Compaction

Strands achieves token efficiency through two default context-management rules:

Tool-result offloading at ~1500 tokens. Long tool outputs (shell logs, search results, test reports) are truncated in the active context; head/tail previews are kept, while full content is written to a stash and retrieved on demand via retrieve_context. This avoids repeatedly paying for the same tokens in every subsequent model call.

Compaction at 85% context-window utilization. When usage hits 0.85, older messages are summarized while the most recent N messages (e.g., 4) are preserved. If overflow still occurs, a context-recovery mechanism continues execution inside the agent loop instead of aborting with a context-length error.

Example: a tool returning 4,000 tokens per call over 20 rounds would accumulate 80,000 tokens of tool output. Offloading removes the recurring cost of re-reading that output in every future inference step.

Strands Context Management documentation showing offloader and auto/agentic strategies
Strands Context Management documentation showing offloader and auto/agentic strategies

Layered Storage: Context Window, Session, Long-Term Memory

Strands separates state into three tiers, mirroring traditional storage hierarchy:

Context Window ≈ working memory – actively constrained via offloading and compaction.

Session ≈ current task state – persisted to ./.agent/sessions; restores full conversation on resume.

Long-term Memory ≈ durable knowledge – managed by MemoryManager with Recall (search), Injection (prepend to prompt), and Extraction (write new facts). Recall and Injection are enabled by default; Extraction requires explicit opt-in.

This prevents dumping all history, preferences, and tool outputs into the context window, reducing cost, latency, and attention dilution.

Strands Sessions vs Memory comparison diagram
Strands Sessions vs Memory comparison diagram

Progressive Skill Disclosure and Default Harness Assembly

Skills use progressive disclosure : only skill names and descriptions enter the system prompt initially; full instructions are loaded via tool call when the model decides they are needed. This mirrors the tool-result offloading philosophy – don't pre-load everything that might be used. create_harness() ships with a pre-assembled stack: shell, read, write, edit, web_fetch, web_search, programmatic_tool_caller, subagent, todos, environment plugins, long-term memory, prompt caching, context manager, and MCP server support. A generalist subagent handles open-ended subtasks while the main agent tracks progress via todos. Even the system prompt is part of the harness defaults, enforcing exploration-before-modification, confirmation for irreversible actions, and re-verification before task completion.

Strands Load Agent Skills documentation showing progressive disclosure
Strands Load Agent Skills documentation showing progressive disclosure

Prompt Caching: Cost Is No Longer Just Token Price

Strands enables caching: auto, leveraging provider-level prompt caching (Bedrock, Anthropic) for system prompts, tool definitions, and stable message prefixes. Because caching capabilities differ across providers, the model interface cannot fully abstract away underlying API differences. The effective cost formula becomes:

Effective Context × Call Count – Cache Hits – Offloaded History + Compaction Cost + Subagent/Tool Call Cost

Identical model pricing can yield wildly different bills depending on harness implementation – hence the need for harness-level benchmarks.

Evaluation Must Shift to Model × Harness × Workload

Strands standardizes the model provider interface (streaming, tool calling, structured output) across Bedrock, Anthropic, OpenAI, Google, Ollama, LiteLLM, Mistral, Llama API, llama.cpp, Vercel, etc. The same harness can run on any Linux-container platform (Cloudflare Containers, Azure Container Apps, Google Cloud Run, Amazon ECS, Bedrock AgentCore, Modal). This decouples both model-from-harness and harness-from-cloud.

However, provider capabilities still differ (prompt caching, tool schema support, reasoning effort, context window, parallel tool calls). Swapping models without rewriting the harness does not guarantee unchanged performance. AWS's own benchmark demonstrates that Model × Harness is a combined variable : the harness controls what the model sees, retains, and when it calls tools; the model determines how well it uses that information. Future evaluation should target Model × Harness × Workload – measuring cost, success rate, and step count on real workloads rather than standalone model leaderboards.

Conclusion

Strands Harness represents a shift in agent scaling: from model-centric improvements to runtime-system engineering. The model sets the capability ceiling, but context, memory, tools, skills, cache, and execution loops determine how much of that ceiling is realized. AWS aims to fix the "body" so developers only swap the "brain." Caveats remain: the benchmark is self-reported, the paper is unpublished, and DeepSeek Harness shows lower cost can come with lower accuracy. Independent verification is needed, but the architectural direction is clear.

References

Strands Agents, Introducing Strands harness: frontier performance with 28% lower token cost , 2026-09-21.

Strands Agents Documentation – Harness, Configuration, Context Management, Memory, Skills.

The New Stack, AWS open-sources an AI agent it says is 45% cheaper than Claude Code and Codex , 2026-09-21.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agentsAWSBenchmarkContext ManagementPrompt CachingAgent RuntimeStrands HarnessModel-Harness Decoupling
DataFunTalk
Written by

DataFunTalk

Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.