Frozen Model Weights, Evolving Agents: ModularRSI Enables Harness Self-Improvement

ModularRSI freezes base model weights and evolves the agent's harness — system mechanisms like loop control, observation handling, tool use, context management, and task completion detection — through execution experience, achieving a 4.86-point gain on Terminal-Bench 2.0 that transfers across tasks, domains, and different base models.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Frozen Model Weights, Evolving Agents: ModularRSI Enables Harness Self-Improvement

IQuest Research, with Beihang University and the University of Manchester, introduces ModularRSI, a framework for Harness Recursive Self-Improvement (Harness RSI). The core idea: keep the base model (M) frozen and evolve the Harness (H) — the surrounding runtime mechanisms that determine what the model sees, how it maintains context, when it calls tools, how it recovers from failures, and when it declares a task complete.

Agent Decomposition and Modular Harness

The authors formalize an agent as A = (M, H). They decompose H into five relatively independent modules:

Agent Loop : manages the reasoning–action–observation iteration.

Observation Management : processes and filters environment feedback.

Tool Use : handles tool selection, invocation, and parameter construction.

Context Management : maintains history, task constraints, and intermediate results.

Task Completion Detection : judges whether the task is truly finished.

Each module evolves independently from the same initial harness.

Evolutionary Process

After a batch of tasks, the system collects execution trajectories. For each task it runs multiple rollouts and compares outcomes:

If all trajectories succeed, it checks for redundant operations or inefficient interactions.

If both success and failure trajectories exist (same task, same environment), it directly contrasts them to pinpoint which decisions changed the result — this is the highest-value signal.

If all fail, it looks for a previously successful version or analyzes dead loops, wrong tool calls, weak error recovery, or premature termination.

Recurring defects across multiple tasks and trajectories are abstracted into structured findings (which harness function is implicated, evidence, proposed modification). Only systemic defects receive high modification priority.

Before a change is written back, it must pass three validation gates:

Program Check : syntax, imports, interface contracts.

Diff Review : rejects task-specific constants, single-instance fixes, or heuristics that only work for the current batch.

Execution Validation : re-runs the modified harness on tasks; any new runtime error causes rollback.

After independent module evolution, a Cross-Module Integration epoch merges improvements. Two additional mechanisms operate: Function Merge consolidates redundant or conflicting functions, and Task-Aware Function Composition dynamically selects and combines functions per task.

Experimental Results

In-Domain Improvement (Terminal-Bench 2.0)

Base harness accuracy: 47.57% → Evolved harness: 52.43% (+4.86 pp).

Single-module contributions: Agent Loop +2.99 pp; Observation Management reduces average interaction steps by ~35%.

Joint all-module evolution degrades performance to 44.19%, revealing strong inter-module coupling. Independent evolution followed by integration outperforms joint evolution.

Cross-Domain Transfer

Harness evolved on Terminal-Bench tasks improves SWE-Bench Verified by 2.40 pp .

Harness evolved on SWE-Bench tasks improves Terminal-Bench 2.0 by 1.83 pp .

Both directions show positive transfer; in-domain transfer yields larger gains.

Cross-Model Transfer

Harness evolved with DeepSeek-V4-Flash Preview, then frozen and paired with other models on Terminal-Bench 2.0:

GLM-5.2: 59.55% → 61.80% (+2.25 pp)

MiniMax-2.5: 41.57% → 44.94% (+3.37 pp)

DeepSeek-V4-Flash: 47.57% → 52.43% (+4.86 pp)

This demonstrates that harness improvements are not tightly bound to a single base model.

Difficulty Distribution Matters

Two evolution data distributions were tested:

Medium-centered (~50% success rate): evolution set score rises from 58.4% to 65.0% (+6.6 pp).

Hard & Easy (mostly trivial or near-impossible tasks): only +0.6 pp.

Reason: ModularRSI learns by contrasting success vs. failure trajectories on the same task. Too-easy tasks yield only successes; too-hard tasks yield only failures. The informative region is where both outcomes occur.

Case Study: Agent Loop Evolution

A concrete trajectory shows stepwise abstraction:

Agent repeats the same invalid command 8 times, then falsely claims completion. System adds guarded_completion to block unsubstantiated completion claims and detect repeated commands.

Later format errors appear; system adds parse_error_recovery to recognize and correct malformed outputs.

Agent then stalls on read-only commands ( ls, cat, grep) without making progress. System extends detection to identify "long exploration without edit/build/test".

Harness evolves planning_checklist to track completed vs. pending requirements, integrating with earlier error-recovery logic.

Function Merge combines guarded_completion, completion_integrity_guard, and planning_checklist into planning_with_guard.

This illustrates how a specific failure is abstracted into general mechanisms (repetition detection, completion judgment, error recovery, stall detection, planning management) — the key distinction from simple experience memorization.

Module Conflict Resolution

Evolved Agent Loop and Verification modules each contained their own task-completion judgment, creating dual completion gates. The agent would pass one gate, get blocked by the other, and loop until timeout. ModularRSI resolved this by keeping Verification as the sole final completion judge while preserving Agent Loop's error-recovery logic.

Behavioral Evaluation Consistency

An LLM judge scored trajectories across dimensions (task validity, interaction quality, reasoning reliability, context/state management, efficiency, robustness). Behavioral scores rose consistently with Terminal-Bench accuracy, confirming that performance gains correspond to observable execution-behavior improvements.

Limitations and Future Work

Stable co-evolution of modules.

Optimal data composition for Harness RSI (difficulty distribution, trajectory diversity, success/failure ratio).

Scaling to larger datasets and more evolution rounds.

Influence of different base model types on harness evolution trajectories.

The work demonstrates a clear phenomenon: with model weights frozen, agents can still discover systemic defects from execution experience, rewrite those insights into the harness, and make the next round of execution build on the previous one — opening a new evolutionary axis for agent self-improvement.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Agent ArchitectureSWE-Benchrecursive self-improvementTerminal-BenchHarness RSIAgent Self-ImprovementModularRSIFrozen Model Weights
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.