MemoHarness: Agent Evolution Shifts from Model Weights to External Harness

MemoHarness introduces a six-dimensional external harness system that adapts agent behavior through case-based experience from execution trajectories, showing improvements on terminal, coding, and finance tasks while acknowledging limited experimental scale and selective cross-task transferability.

DataFunTalk
DataFunTalk
DataFunTalk
MemoHarness: Agent Evolution Shifts from Model Weights to External Harness

The article analyzes the MemoHarness paper (arXiv:2607.14159), which proposes shifting agent evolution from internal model parameters to an external control layer called the Agent Harness. The harness comprises six editable dimensions: D1 Context Assembly (organizing instructions, constraints, retrieved materials, examples), D2 Tool Interaction (deciding when to call tools, how much to retrieve, whether to rerank evidence), D3 Generation Control (adjusting max output, temperature, candidate sampling), D4 Task Orchestration (choosing between single-call and multi-stage workflows like plan–execute–correct), D5 Memory Management (retaining key state, summarizing trajectories, removing stale context), and D6 Output Processing (extracting answers, verifying format, fallback strategies on failure).

Unlike prompt engineering, which focuses on what instructions the model receives, harness engineering concerns the entire system in which the model operates. A fixed global harness mismatches diverse tasks: simple coding may need one call, while long-horizon terminal tasks require continuous environment reading, state saving, exception handling, and result verification.

MemoHarness operates in two phases. In the development/search phase (termed "training-time optimization" though model weights stay frozen), the system starts from a minimal harness (no examples, no structured scaffolding, no external tools, deterministic single call, no cross-call memory, raw output). It searches harness configurations on tasks with reference answers, ranking candidates by correctness first, token cost second. Each execution records task features, current harness, delta from previous config, full model/tool trajectory, score, token usage, and failure-dimension diagnosis — forming case-level experience. Periodically, global patterns are distilled across cases into a two-layer experience base.

In the test phase, the experience base is frozen. For a new unlabeled task, the system retrieves similar success/failure cases and global patterns, then adapts the global harness into a task-specific configuration — a single case-based adaptation without labels, feedback, or further search.

Experiments on three benchmarks show gains: Terminal-Bench 0.722→0.806, LiveCodeBench 0.900→0.967, FinanceAgent 0.600→0.767. Improvement is larger on tool-intensive, long-horizon tasks (FinanceAgent) than on near-saturated coding (LiveCodeBench). Cross-dataset transfer is selective: Terminal-Bench harness improves MMMLU, StrongReject, SWE-Bench Pro, but not HumanEvalFix or Reasoning-Gym-Easy; LawBench is mixed. Cross-model transfer (harness searched with GPT-5.3-Codex applied to six other models) yields average +0.098 success rate (range +0.038 to +0.233), suggesting some strategies are model-agnostic but not conclusively. Cost on Terminal-Bench 18 held-out tasks: 14.18M input tokens (13.32M cached), $6.89 at public pricing — lower than Codex and Claude Code but higher than OpenCode; advantage depends on high cache reuse.

The article emphasizes limitations acknowledged by the authors: point estimates without confidence intervals or significance tests; baselines not strictly same-model harness ablations; incomplete ablations of experience base, global patterns, and case-based adaptation; no online accumulation during deployment (each test task gets one adaptation, results not fed back). A concurrent paper (arXiv:2607.12227) raises broader evaluation standards: harness evolution must be compared against budget-matched parallel sampling and sequential correction, with strict separation of search and final evaluation tasks.

The core contribution is reframing agent evolution: instead of updating model weights, improvements manifest as adding tools, changing retrieval, compressing context, adjusting workflows, or adding output validators. Each harness change is versionable, testable, rollbackable, and auditable. Organizations can accumulate proprietary execution trajectories, error diagnoses, eval suites, and harness versions — creating defensible capability beyond the base model. The model becomes a swappable reasoning engine; the hard-to-replicate asset moves to the external "operating system."

Agent evolution may not happen in the model. At minimum, the model is no longer the only evolvable component.

References: [1] Yue Huang et al. "MemoHarness: Agent Harnesses That Learn from Experience." arXiv:2607.14159 (2026-07-14). [2] MemoHarness code repo: https://github.com/HowieHwong/MemoHarness. [3] Yike Wang et al. "Rethinking the Evaluation of Harness Evolution for Agents." arXiv:2607.12227 (2026-07-14).

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LLM agentsHarness EngineeringAgent HarnessTerminal-BenchLiveCodeBenchMemoHarnesscase-based adaptationexecution trajectoriesFinanceAgentexternal control system
DataFunTalk
Written by

DataFunTalk

Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.