ScienceBuddy: Open-Source AI Research Agent with Recursive Self-Improvement

ScienceBuddy is an open-source AI research agent that uses a recursive dual-loop architecture to continuously improve its harness and model weights from real user interactions, achieving 42.2% to 73.3% accuracy gains on scientific benchmarks across genomics, molecular biology, and pharmacology domains.

PaperAgent
PaperAgent
PaperAgent
ScienceBuddy: Open-Source AI Research Agent with Recursive Self-Improvement

Introduction

ScienceBuddy is the academic counterpart of Tencent's WorkBuddy AI office agent, now open-sourced. It functions as a self-improving scientific workbench that turns everyday researcher interactions — requests, feedback, result reviews — into training tasks and rubrics, driving two coupled improvement loops: harness adaptation and model reinforcement learning.

Core Mechanism: Recursive-in-Recursive Dual Loop

The system architecture is a "double loop" (Figure 2) with three layers:

Inner recursion (model fixed) : A fixed assistant model (GPT-6 Astra) diagnoses failed trajectories and proposes bounded modifications to harness instructions, skills, and context settings. Candidates are accepted only if they strictly improve on paired development evaluations; otherwise they are rolled back.

Outer recursion (harness fixed) : The selected harness calibrates task difficulty, runs fresh on-policy trajectories, uses rubric scores as rewards, and updates model weights via GRPO.

Coupled rolling : The new model paired with the inherited harness is re-evaluated and deployed without interrupting online service. Researchers continue using the system, and their new interactions feed the next round.

Figure 2|Recursive-in-Recursive framework
Figure 2|Recursive-in-Recursive framework

From Collaboration to Training Tasks

Real collaboration sessions are packaged into self-contained task bundles (instruction, assets, environment, rubric test scripts). The same bundle serves both rejection-sampling SFT and on-policy RL. The paper emphasizes that researcher approval is not treated as ground-truth scientific fact; rubrics undergo conflict resolution before entering the task bank.

Figure 3|From collaboration to Harbor task
Figure 3|From collaboration to Harbor task

Product Interface

ScienceBuddy is a deployed workbench, not just a training framework:

224 tools across 22 functional modules covering genomics, molecular & cancer biology, pharmacology, bioimaging, literature search, and database querying.

Runtime supports Python, R, and Bash with a built-in local data lake.

Documents, tables, sequences, and images can be dragged directly into the chat.

Full auditability: the Trajectory panel exposes every tool call's input, output, and latency.

Figure 4|Multimodal analysis workbench
Figure 4|Multimodal analysis workbench
Figure 6|Execution trajectory inspectable
Figure 6|Execution trajectory inspectable

Turning Interaction into Training Signals

Two real-world cases illustrate the pipeline: a small-cell lung cancer JAK1 research design and an ARL4C paper presentation. A researcher's utterance like "focus on JAK1 first, mechanism later" is translated into a task goal, rubric, and deliverable requirements (Figure 7). This step is the paradigm's bottleneck — collaboration quality directly determines training signal quality.

Figure 7|From request to task specification
Figure 7|From request to task specification

Experiments: Numbers Rising

Evaluation uses 895 tasks from LAB-Bench and Biomni-Eval1 spanning four families: literature reading, database judgment, experiment debugging, and gene variant assessment (Table S1).

Table S1|Task composition
Table S1|Task composition

Three rounds of co-evolution (starting from Qwen3.5-4B, 10 harness search steps + 20 RL updates per round):

Harness validation accuracy: 38.9%→44.4%, 34.4%→46.7%, 61.1%→70.0% across rounds.

Test-set single-pass accuracy: 42.2% → 73.3% , with 33.3% of questions flipping from wrong to right and only 2.2% regressing.

All four task families improved.

Figure 8|Three-round co-evolution dynamics
Figure 8|Three-round co-evolution dynamics

Harness-only evolution (model weights frozen): after 24 adaptive batches, validation first-pass accuracy rose 31.1% → 51.1% (+20 pp) . The final harness accumulated 4 instructions and 9 scoped skills (e.g., gene-set membership check, cytoband query).

Figure 9|Harness adaptation
Figure 9|Harness adaptation

Model-only RL (harness frozen): after two hours of training, pass@4 problem coverage increased 48.3% → 67.8% (+19.5 pp) — more problems solvable within the same budget.

Figure 10|Model learning under fixed harness
Figure 10|Model learning under fixed harness

Conclusion

The paper's main contribution is closing the loop between product interaction and self-improvement : more usage enriches the task bank and rubrics, which strengthen both harness and model, enabling harder tasks — a concrete instantiation of Silver and Sutton's "Era of Experience" vision.

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
Arxiv: 2609.17523
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Open sourceAI AgentReinforcement Learningscientific researchrecursive self-improvementScienceBuddyharness adaptationLAB-Bench
PaperAgent
Written by

PaperAgent

Daily updates, analyzing cutting-edge AI research papers

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.