ModularRSI: Agents Self-Improve by Evolving Harness, Not Model Weights

ModularRSI demonstrates that AI agents can continuously self-improve by evolving their harness—the surrounding system mechanisms like tool use, context management, and task completion detection—while keeping the base model frozen, achieving performance gains on Terminal-Bench and SWE-Bench that transfer across domains and models.

Machine Heart
Machine Heart
Machine Heart
ModularRSI: Agents Self-Improve by Evolving Harness, Not Model Weights

Introduction

IQuest Research, Beihang University, and the University of Manchester introduced ModularRSI , a framework that explores Harness Recursive Self-Improvement (Harness RSI) . The core question: when the base model (M) is frozen, can an agent (A = (M, H)) continuously improve by modifying its Harness (H) —the system mechanisms that govern how the model interacts with the environment?

The Harness includes what environment information the model sees, how context is maintained, when tools are called, how failures are recovered, and when a task is declared complete. ModularRSI shows that evolving these mechanisms yields measurable, transferable gains without any model weight updates.

Method: Modular Harness Decomposition and Evolution

ModularRSI decomposes the Harness into five relatively independent modules:

Agent Loop : manages the reasoning–action–observation iteration.

Observation Management : processes and filters environment feedback.

Tool Use : handles tool selection, invocation, and parameter construction.

Context Management : maintains history, task constraints, and intermediate results.

Task Completion Detection : judges whether a task is truly finished.

Each module evolves independently from the same initial Harness. After a batch of tasks, the system collects execution trajectories and performs multiple rollouts per task. It then compares trajectories in three scenarios:

All trajectories succeed → check for redundant operations or inefficient interactions.

Mixed success and failure → directly contrast trajectories to identify decisive mechanism differences.

All fail → search for historical successes or analyze dead loops, wrong tool calls, weak recovery, or premature termination.

Findings are structured as which Harness function is problematic, evidence from trajectories, and suggested modification . Only defects recurring across multiple tasks and trajectories receive high modification priority.

Before code is written back, it passes three validation gates:

Program Check : syntax, imports, interface contracts.

Diff Review : rejects task-specific constants, single-instance solutions, or heuristics that only work for the current batch.

Execution Validation : re-runs the modified Harness on tasks; runtime errors cause rollback.

After independent evolution, a Cross-Module Integration epoch merges improvements. Two additional mechanisms reduce redundancy and conflict: Function Merge combines similar or redundant functions, and Task-Aware Function Composition dynamically selects and composes functions per task.

ModularRSI framework overview: five Harness modules, trajectory analysis, Harness Evolution, and Validation Gate
ModularRSI framework overview: five Harness modules, trajectory analysis, Harness Evolution, and Validation Gate

Experimental Setup and Main Results

The team built a pool of 2,000 high-quality instances, sampling two non-overlapping evolution sets: 120 Terminal-Bench-related and 120 SWE-Bench-related instances. Evolution and model selection never used final benchmark data. Base model: DeepSeek-V4-Flash Preview.

Terminal-Bench 2.0 accuracy rose from 47.57 to 52.43 (+4.86 percentage points). Module-level ablation showed:

Agent Loop contributed the largest accuracy gain: +2.99 pp .

Observation Management reduced average interaction steps by ~35% .

Jointly evolving all modules together ( Joint All-Module Evolution ) degraded performance to 44.19 , revealing strong inter-module coupling. The modular evolve-then-integrate approach avoided this.

Transferability: Across Domains and Models

Cross-domain: A Harness evolved on Terminal-Bench tasks improved SWE-Bench Verified by +2.40 pp . Conversely, a SWE-evolved Harness improved Terminal-Bench 2.0 by +1.83 pp . In-domain gains were larger, but both cross-domain directions showed positive transfer.

Cross-model: The Harness evolved with DeepSeek-V4-Flash was frozen and paired with other models on Terminal-Bench 2.0: GLM-5.2: 59.55 → 61.80 (+2.25 pp) MiniMax-2.5: 41.57 → 44.94 (+3.37 pp) DeepSeek-V4-Flash: 47.57 → 52.43 (+4.86 pp)

This indicates that part of the Harness-layer experience is not tightly bound to a specific base model. Mechanisms like reducing useless loops, improving error recovery, and making completion judgments more reliable generalize across models.

Case study: evolution from guarded_completion, completion_integrity_guard, planning_checklist to planning_with_guard
Case study: evolution from guarded_completion, completion_integrity_guard, planning_checklist to planning_with_guard

Data Difficulty Determines Learning Efficiency

Two data distributions were tested:

Medium-centered : ~50% success rate (40–60%). Evolution set score rose from 58.4% to 65.0% (+6.6 pp).

Hard & Easy : mostly very easy or very hard tasks. Gain was only +0.6 pp .

ModularRSI learns by contrasting success and failure trajectories on the same task. Tasks that are too easy yield only successes; too hard yield only failures. The "sometimes succeed, sometimes fail" region provides the highest information density for extracting generalizable mechanisms.

Case Study: Agent Loop Evolution

A concrete evolution trace illustrates how specific failures are abstracted into general mechanisms:

Agent executed the same invalid command 8 times, then repeatedly claimed completion. Verifier rejected until budget exhausted. → Added guarded_completion to block unsubstantiated completion claims and detect repeated commands.

Later trajectories showed format errors. → Added parse_error_recovery to recognize and correct malformed output.

Agent then stalled on long sequences of read-only commands ( ls, cat, grep) without editing, building, or testing. → Extended mechanism to detect "prolonged exploration without progress".

Harness evolved planning_checklist to track completed vs. pending requirements, integrated with earlier recovery logic.

Function Merge combined these into planning_with_guard.

This shows Harness RSI abstracts concrete failures into reusable mechanisms—distinct from simple experience memorization.

Module Conflict and Resolution

Independently evolved Agent Loop and Verification modules each contained their own task-completion judgment. When combined, the dual completion gates caused the agent to loop between "complete → verify → re-complete" until timeout. The fix: retain Verification as the sole final gate while preserving Agent Loop's error-recovery logic. This underscores the need for explicit inter-module optimization.

Behavioral Evaluation Consistency

An LLM judge scored trajectories on task validity, interaction quality, reasoning reliability, context/state management, efficiency, and robustness across evolution generations. Behavioral scores correlated with Terminal-Bench accuracy, confirming that performance gains reflect observable execution improvements.

Conclusion and Future Work

ModularRSI provides a concrete engineering implementation of Harness RSI. While model scaling expands capability boundaries, Harness optimization improves how those capabilities are deployed in real execution—tool use, context management, error recovery, efficiency, and completion judgment. As agents tackle longer, more complex tasks, both axes will be complementary.

Open questions remain: stable co-evolution of modules, optimal data for Harness RSI, scaling to more data and generations, and how different base models shape Harness evolution trajectories. The paper, code, and blog are available at:

arXiv: https://arxiv.org/abs/2609.14857 Code: https://github.com/IQuestLab/ModularRSI Blog:

https://recursive-self-improvement.notion.site/blog-1-modularrsi-toward-generalizable-harness-rsi
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Cross-Domain TransferSWE-BenchTerminal-BenchCross-Model TransferHarness RSIAgent Self-Improvementfrozen modelModularRSI
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.