Frozen Model Weights, Evolving Agents: ModularRSI Enables Harness Self-Improvement
ModularRSI freezes base model weights and evolves the agent's harness — system mechanisms like loop control, observation handling, tool use, context management, and task completion detection — through execution experience, achieving a 4.86-point gain on Terminal-Bench 2.0 that transfers across tasks, domains, and different base models.
IQuest Research, with Beihang University and the University of Manchester, introduces ModularRSI, a framework for Harness Recursive Self-Improvement (Harness RSI). The core idea: keep the base model (M) frozen and evolve the Harness (H) — the surrounding runtime mechanisms that determine what the model sees, how it maintains context, when it calls tools, how it recovers from failures, and when it declares a task complete.
Agent Decomposition and Modular Harness
The authors formalize an agent as A = (M, H). They decompose H into five relatively independent modules:
Agent Loop : manages the reasoning–action–observation iteration.
Observation Management : processes and filters environment feedback.
Tool Use : handles tool selection, invocation, and parameter construction.
Context Management : maintains history, task constraints, and intermediate results.
Task Completion Detection : judges whether the task is truly finished.
Each module evolves independently from the same initial harness.
Evolutionary Process
After a batch of tasks, the system collects execution trajectories. For each task it runs multiple rollouts and compares outcomes:
If all trajectories succeed, it checks for redundant operations or inefficient interactions.
If both success and failure trajectories exist (same task, same environment), it directly contrasts them to pinpoint which decisions changed the result — this is the highest-value signal.
If all fail, it looks for a previously successful version or analyzes dead loops, wrong tool calls, weak error recovery, or premature termination.
Recurring defects across multiple tasks and trajectories are abstracted into structured findings (which harness function is implicated, evidence, proposed modification). Only systemic defects receive high modification priority.
Before a change is written back, it must pass three validation gates:
Program Check : syntax, imports, interface contracts.
Diff Review : rejects task-specific constants, single-instance fixes, or heuristics that only work for the current batch.
Execution Validation : re-runs the modified harness on tasks; any new runtime error causes rollback.
After independent module evolution, a Cross-Module Integration epoch merges improvements. Two additional mechanisms operate: Function Merge consolidates redundant or conflicting functions, and Task-Aware Function Composition dynamically selects and combines functions per task.
Experimental Results
In-Domain Improvement (Terminal-Bench 2.0)
Base harness accuracy: 47.57% → Evolved harness: 52.43% (+4.86 pp).
Single-module contributions: Agent Loop +2.99 pp; Observation Management reduces average interaction steps by ~35%.
Joint all-module evolution degrades performance to 44.19%, revealing strong inter-module coupling. Independent evolution followed by integration outperforms joint evolution.
Cross-Domain Transfer
Harness evolved on Terminal-Bench tasks improves SWE-Bench Verified by 2.40 pp .
Harness evolved on SWE-Bench tasks improves Terminal-Bench 2.0 by 1.83 pp .
Both directions show positive transfer; in-domain transfer yields larger gains.
Cross-Model Transfer
Harness evolved with DeepSeek-V4-Flash Preview, then frozen and paired with other models on Terminal-Bench 2.0:
GLM-5.2: 59.55% → 61.80% (+2.25 pp)
MiniMax-2.5: 41.57% → 44.94% (+3.37 pp)
DeepSeek-V4-Flash: 47.57% → 52.43% (+4.86 pp)
This demonstrates that harness improvements are not tightly bound to a single base model.
Difficulty Distribution Matters
Two evolution data distributions were tested:
Medium-centered (~50% success rate): evolution set score rises from 58.4% to 65.0% (+6.6 pp).
Hard & Easy (mostly trivial or near-impossible tasks): only +0.6 pp.
Reason: ModularRSI learns by contrasting success vs. failure trajectories on the same task. Too-easy tasks yield only successes; too-hard tasks yield only failures. The informative region is where both outcomes occur.
Case Study: Agent Loop Evolution
A concrete trajectory shows stepwise abstraction:
Agent repeats the same invalid command 8 times, then falsely claims completion. System adds guarded_completion to block unsubstantiated completion claims and detect repeated commands.
Later format errors appear; system adds parse_error_recovery to recognize and correct malformed outputs.
Agent then stalls on read-only commands ( ls, cat, grep) without making progress. System extends detection to identify "long exploration without edit/build/test".
Harness evolves planning_checklist to track completed vs. pending requirements, integrating with earlier error-recovery logic.
Function Merge combines guarded_completion, completion_integrity_guard, and planning_checklist into planning_with_guard.
This illustrates how a specific failure is abstracted into general mechanisms (repetition detection, completion judgment, error recovery, stall detection, planning management) — the key distinction from simple experience memorization.
Module Conflict Resolution
Evolved Agent Loop and Verification modules each contained their own task-completion judgment, creating dual completion gates. The agent would pass one gate, get blocked by the other, and loop until timeout. ModularRSI resolved this by keeping Verification as the sole final completion judge while preserving Agent Loop's error-recovery logic.
Behavioral Evaluation Consistency
An LLM judge scored trajectories across dimensions (task validity, interaction quality, reasoning reliability, context/state management, efficiency, robustness). Behavioral scores rose consistently with Terminal-Bench accuracy, confirming that performance gains correspond to observable execution-behavior improvements.
Limitations and Future Work
Stable co-evolution of modules.
Optimal data composition for Harness RSI (difficulty distribution, trajectory diversity, success/failure ratio).
Scaling to larger datasets and more evolution rounds.
Influence of different base model types on harness evolution trajectories.
The work demonstrates a clear phenomenon: with model weights frozen, agents can still discover systemic defects from execution experience, rewrite those insights into the harness, and make the next round of execution build on the previous one — opening a new evolutionary axis for agent self-improvement.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
