Dream-RSI: How Agents 'Dream' on Past Searches to Optimize Exploration Budgets

Dream-RSI introduces a recursive self-improvement framework where AI agents improve exploration strategies by replaying historical search trees offline, enabling budget allocation decisions without retraining models, demonstrated across algorithm engineering, math optimization, and GPU kernel tasks with reduced search costs.

Architect
Architect
Architect
Dream-RSI: How Agents 'Dream' on Past Searches to Optimize Exploration Budgets

Agent Dreams: What Dream-RSI Optimizes

When agents perform complex search, they typically only work "during the day": generating candidates, running evaluations, and iterating. Dream-RSI adds a "nighttime review" phase. After each search round, the system saves the full discovery tree and lets a new exploration policy first test-run on historical paths before deciding how to allocate the next round's budget. This "dreaming" is not imagination—it re-examines paths the agent has already taken.

The key insight: Dream-RSI does not modify model parameters; it improves how the agent finds answers.

What the Agent Dreams About

In complex search, the hard part is not generating more candidates but deciding where to spend the next compute budget: deepen the current best branch, explore a new direction, open more parallel branches, or abandon a stalled branch. Hard-coding these decisions makes the strategy brittle when the search space changes. Online testing of a new strategy is slow because feedback arrives only after a full discovery round completes.

Dream-RSI targets this layer. The paper, from Google, Google DeepMind, University of Maryland, and University of Virginia, extracts node selection, branching, concurrency, depth, and stopping decisions into a separate exploration policy code that can be continuously improved. The Coding Agent still proposes candidates; Dream-RSI adjusts how candidates are found and how budget is spent .

Daytime Online Search, Nighttime Path Replay

The Coding Agent, Evaluator, initialization, and per-round resource constraints stay fixed. Only the exploration policy code changes. This code schedules search by deciding:

Which node to select next

Whether to open more branches from a node

How to compose parallel batches

Which directions to deepen

When to stop current exploration

In the agent loop, the Coding Agent proposes, the Evaluator scores, and the exploration policy decides where to continue. Previously, scheduling logic was entangled with the agent loop. Dream-RSI isolates it, enabling independent evaluation and rollback.

Dream-RSI 把哪一层拿出来优化
Dream-RSI 把哪一层拿出来优化

Figure 1: Candidates proposed by Coding Agent; exploration policy decides how budget flows along the search tree.

Fixed policies (called Recursive Fixed Exploration) allocate budget uniformly—same number of workspaces, same steps per round. Real search is uneven: early on, broad branching pays off; later, budget should concentrate on promising directions; stalled branches should release resources. Fixed policies spread budget evenly; Dream-RSI tries to make budget follow search evidence.

固定探索与 Dream-RSI 的策略差异
固定探索与 Dream-RSI 的策略差异

Figure 2: Under identical model, evaluator, and per-round budget, the difference is how budget is allocated along the tree.

The controlled baseline uses the same Coding Agent, Evaluator, initialization, and per-round constraints. First-round behavior is identical. From round two onward, the baseline keeps the original policy while Dream-RSI modifies the policy code based on historical replay. This isolates the effect of exploration scheduling from model changes or extra budget.

A Dream in Three Acts

Online Exploration: Leaving Behavioral Traces

The current exploration policy drives the Coding Agent. When a node is selected, its workspace and observations are restored, the Coding Agent generates a new candidate, the Evaluator returns a score and diagnostics, and the new node is attached to the tree. The paper records each node's parent, saved file state, generated artifacts, evaluation diagnostics, and score. After each round, the entire discovery tree enters the history set—preserving the actual paths taken, not just the final answer.

Replay: Re-walking the Historical Tree

After a round, the historical discovery tree is turned into a Replay Simulator. A new policy starts at the root, sees only the already-revealed prefix, then decides which branches to open, which leaves to continue, how to form parallel batches, and when to stop. During replay, selected nodes do not call the Coding Agent or re-run the Evaluator; the system reveals recorded results along parent-child links and lets the policy continue deciding.

Replay optimizes three quantities simultaneously: best score reached, number of generations used, and parallel batch execution efficiency. The simulator is not a predictive world model; it is an empirical replay environment built from real search records, covering only the search space already explored. Replay does not generate new candidates; it only reschedules paths that have already occurred.

Policy Development Agent: Rewriting the Controller

A fixed LLM-based Policy Development Agent reads the candidate policy's replay trajectories and scores, analyzes which choices were effective and which wasted budget, then rewrites the exploration policy code. The modified version is evaluated on the same batch of historical trees. The system selects the version with higher average replay score. The Policy Development Agent compares branch orders, parallel batches, and stopping rules without restarting the full discovery process. The chosen policy is deployed back online, producing new discovery trees for the next replay round.

Dream-RSI 的三段闭环
Dream-RSI 的三段闭环

Figure 3: Replay lowers policy comparison cost; online exploration continues to open new search space.

The loop: online exploration leaves history → history replay compares policies → selected policy returns online to generate new data.

Why a Text Summary Is Not Enough

Dream-RSI relates to Agent Memory but stores behavioral data, not just judgments. A textual summary like "this direction performed poorly, explore others" loses critical details: which branches were tried, their parent-child relationships, which candidates were slow vs. invalid, when budget was consumed, whether a direction was stopped prematurely or fully explored, and whether a different visitation order would have yielded the same result earlier. Behavioral replay preserves the branches, order, costs, and failure evidence—inputs closer to computable decisions for "how to allocate budget next."

The paper compares two history-use methods on ConvDiv: compressing history into directional text injected into prompts vs. keeping history as an interactive replay tree. Under equal budget, explicit directional guidance underperformed the no-guidance version in both exploration paradigms. The explanation: long-horizon search maintains multiple parallel threads; premature direction narrows the search space and reduces diversity. Text summaries tell the model "what happened"; replay compares "how would a different scheduling have visited the already-occurred search." They solve different problems.

Paper Results: Where Gains Come From

Dream-RSI covers 8 tasks across 3 domains: algorithm engineering, mathematical optimization, and GPU kernel engineering. Controls use identical Coding Agent, Evaluator, initialization, and per-round resource constraints; only exploration strategy differs.

Algorithm Engineering (Lasso regularization path): On 6 held-out datasets, measuring downstream runtime. With Gemini-3.1-Pro, fixed exploration used 550 discovery-agent calls (avg 3,587.1 ms); Dream-RSI used 317 calls (avg 2,931.0 ms). With Gemini-3.7-Flash, calls dropped from 3,200 to 1,879 , avg runtime from 2,516.7 ms to 2,350.6 ms. Baselines ran 10 workspaces × 11 steps and 32 workspaces × 20 steps per round; Dream-RSI kept the same per-round budget but changed policy in later rounds.

GPU Kernel Engineering: To reach comparable performance, Dream-RSI required generations at 1/2.43 (VGG16) and 1/1.79 (LayerNorm) of fixed exploration. On ConvDiv and ConvMax, with similar budget, Dream-RSI achieved 2.09× and 1.44× baseline performance. In 3 mathematical optimization tasks, Dream-RSI matched or exceeded the chosen baseline on 2 tasks.

These results support a specific conclusion: in long-horizon, many-branch, high-evaluation-cost discovery tasks, adjusting exploration scheduling can reduce search cost and improve outcomes. The magnitude is context-dependent—not all agent tasks see the same gain, and "50× budget savings" claims from competitor benchmarks do not generalize. The paper discusses that magnitude only in a specific SimpleTES task comparison; versus the controlled fixed exploration, conclusions are narrower.

More importantly, exploration policy is no longer a few fixed lines in a loop; it becomes a separable component that can be experimented with, compared, and rolled back.

How It Connects to the Harness

Previous work decomposed agents into Model, Loop, Harness, and Evaluator. Model generates next steps; Loop sequences actions; Harness manages context, tools, state, permissions, recovery, and observation; Evaluator provides external feedback. Dream-RSI optimizes the exploration control logic inside the Harness.

This aligns with Karpathy's context engineering: agent engineering is not just better prompts but organizing context, tools, state, history, and control flow. Dream-RSI moves one layer up: from "what the model sees next" to "where the whole search process goes next."

In this layered view, Dream-RSI sits between Control and RSI:

Model : What candidates can be proposed next? → Coding Agent

Environment : How do candidates actually perform? → Evaluator / Execution Environment

History : What has already happened? → Discovery tree / execution trace

Control : Where to allocate resources next? → Exploration policy

RSI : How to improve this control logic? → Policy Development Agent + replay

It also connects to Plan Mode. A static pre-work plan is quickly invalidated by new evidence. Dream-RSI makes "how to explore next" an executable policy, then uses historical trajectories to verify whether the policy spends budget in the right places.

What Replay Can and Cannot Compare

Replay is cheap but has no free future. It can compare:

What results occur if a different set of existing branches is visited first

How budget consumption changes with different concurrency batches

Whether stopping a direction earlier frees resources for other branches

How dynamic width vs. fixed width utilizes historical results

It cannot answer:

Whether a never-generated new candidate would succeed

Whether a new policy is reliable on a completely different task distribution

What results would appear if online exploration ventured into new branches

How unrecorded environment changes would affect the policy

Replay 能比较什么,不能替代什么
Replay 能比较什么,不能替代什么

Figure 4: Replay suits filtering historical scheduling; it cannot turn unoccurred search into known facts.

Therefore, two loops are needed:

Offline replay loop:
  compare policies on realized history

Online exploration loop:
  discover new candidates and expand history

Offline replay alone risks overfitting to old trees without capacity for new spaces. Treat replay as a low-cost filter; online exploration as ground truth. Offline filters out clearly bad policies; online results decide whether history needs expansion and whether the policy truly works.

Putting "Dreaming" Back into Agent Architecture

First, Record Search Behavior

To later improve how an agent works, saving only the final code or a summary is insufficient. At minimum, record: the goal, visited nodes, tools invoked, why the next step was chosen, budget spent, evaluation results, and what triggered stopping. The value of these records is not longer logs but enabling subsequent policy comparison.

Version Policy and Executor Separately

Underlying models may upgrade, evaluators may change, exploration policies may evolve. Mixed together, performance changes are hard to attribute. Engineering practice: version separately:

Model version: who proposed candidates

Evaluator version: who judged results

Policy version: how resources were allocated

History version: which experience batch the comparison is based on

Only then can you tell if performance shifts come from model, evaluator, or scheduling. When policy degrades, you can roll back to the previous controller version.

First, Assess Whether the Task Warrants It

Not every agent needs Dream-RSI-style meta-exploration. If tasks are short, candidate space small, evaluation cheap, simply adding more attempts may be simpler. Only when search is long, branches many, evaluation expensive, and fixed policy visibly wastes budget does separate exploration control optimization pay off. Building a full Replay platform for a 10-call search is overkill; for thousands of expensive evaluations, a dedicated policy control plane is justified.

How Far From "Fully Self-Improving"?

Dream-RSI addresses recursive improvement at the exploration layer. It does not let the agent auto-define goals, invent evaluators, or prove unlimited recursive strengthening on any new task. The paper answers a narrower question: how to compare and improve an agent's search strategy without re-paying full online cost each time.

This scope makes it a more deployable engineering component. The system isolates a slice of control logic, gives it inputs, feedback, and stopping conditions, then validates offline-screened policies with online search.

Dream-RSI's loop in four sentences:

探索产生经验
经验变成回放环境
回放比较策略
策略回到线上

"Dreaming" is just a vivid name for the replay phase. The concrete engineering actions: turn history from static logs into executable experience; turn exploration from fixed scripts into verifiable controllers. What the controller learns depends on what the history tree records; what replay can filter depends on whether online exploration ever traversed relevant branches. The earlier engineering questions: which search decisions are worth recording, which feedback suffices to compare policies, and when must we return online for validation.

References

Dream-RSI: Recursive Self-Improvement through Evolving Worlds (https://github.com/zhengkid/Dream-RSI)

Dream-RSI Official Project Page (https://www.dream-rsi.com/)

arXiv:2609.14858 (https://arxiv.org/abs/2609.14858)

Dream-RSI Paper PDF (https://github.com/zhengkid/Dream-RSI/blob/main/papers/Dream-RSI.pdf)

Andrej Karpathy, public discussion on context engineering (https://x.com/karpathy)

Simon Willison, Context engineering (https://simonwillison.net/2025/Jun/27/context-engineering/)

From One LLM Call to Complete Harness: System Understanding of How Agents Work (https://mp.weixin.qq.com/s?__biz=MzAwNjQwNzU2NQ==&mid=2650410533&idx=1&sn=dadafe2c6a74514249528c486bd4ccf5&scene=21#wechat_redirect)

Plan Mode's Death? Agents Need a Verifiable Work Loop (https://mp.weixin.qq.com/s?__biz=MzAwNjQwNzU2NQ==&mid=2650410550&idx=1&sn=190b52593964530f7757d8c26d172882&scene=21#wechat_redirect)

Agent Memory (https://mp.weixin.qq.com/s?__biz=MzAwNjQwNzU2NQ==&mid=2650409319&idx=1&sn=0301eb71a592e20188e071a238fcd9c8&scene=21#wechat_redirect) and Recursive Self-Improvement (https://mp.weixin.qq.com/s?__biz=MzAwNjQwNzU2NQ==&mid=2650410491&idx=1&sn=df0ac978cc6980fb0c6d8e1d57abccb9&scene=21#wechat_redirect) series

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI researchLLM AgentsGoogle DeepMindrecursive self-improvementagent explorationDream-RSIreplay simulatorsearch budget optimization
Architect
Written by

Architect

Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.