Raven V0.2.0: Harness of Harnesses Orchestrates Agents & Enables Self-Improvement
EverMind's Raven V0.2.0 introduces a Harness of Harnesses architecture that orchestrates specialized agents like Raven-Research, Raven-Code, Raven-Design, Raven-Oncall, Claude Code, and Codex via DAG-based task graphs, while its experimental Curator component enables runtime Harness self-modification across Memory, Planning, Capability, and Action policies, demonstrating superior benchmark results in multi-agent orchestration, deep research, coding, design, and continuous execution tasks.
Recursive Self-Improvement (RSI) is often associated with model parameter updates, but the article argues — drawing on the Complementary Learning Systems theory (Kumaran, Hassabis & McClelland, 2016) — that the hippocampus rapidly encodes episodic experience while the cortex slowly consolidates reusable skills. By analogy, the LLM corresponds to the cortex (slow consolidation), whereas the Harness (memory, skills, prompts, control code, decision policies) should act like the hippocampus, adapting quickly. EverMind's Raven V0.2.0 operationalizes this insight through two core designs.
Two Core Designs
1. Harness Designed for RSI. Raven decomposes the Harness into four independently replaceable categories: Modules (playbooks and sub-Harness compositions), Code (execution-time policy code such as pre-checks), Prompt (system prompts and onboarding manuals), and Policy (tool exposure and gate configuration). An experimental runtime Curator reads the currently assembled Harness, task requirements, execution traces, and user feedback, then generates concrete changes to any of the four categories. Changes undergo declaration checks, real assembly, and a pre-flight run; only after passing are they installed, with automatic rollback on failure. All modifications target Raven's native extension points — no core source changes required.
2. The Harness of Harnesses. Raven orchestrates four deeply optimized native agents — Raven-Research, Raven-Code, Raven-Design, and Raven-Oncall — alongside third-party professional agents such as Claude Code and Codex via ACP, CLI, or OpenAI-compatible APIs. Raven handles member selection, task decomposition, dependency scheduling, and result handoff, while EverOS preserves cross-session context and experience, enabling cross-model, cross-framework capability composition.
Harness of Harnesses: Organizing Professional Agents
Complex tasks require deciding what to research first, which agent should act, what can run in parallel, how results are handed off, and how to adjust when issues arise. Raven's outer layer manages members, tasks, and dependencies; each member's internal Harness governs how information is used, tools are called, and actions proceed. The four native agents cover research, coding, design, and continuous execution respectively. Third-party agents retain their own Harness and execution mechanics.
From Goal to Task Graph
Raven translates a high-level goal into a Directed Acyclic Graph (DAG) of tasks. Nodes without mutual dependencies launch in parallel; dependent nodes wait for upstream completion. On node start, Raven assembles the task objective, role responsibilities, and explicit references to upstream outputs or artifact entry points. If a predecessor fails, its dependents are skipped while unrelated branches continue. The resulting execution chain can be saved as a Playbook for reuse in similar scenarios.
Example: a cantilever beam simulation task splits into research (Raven-Research), implementation (Raven-Code), and experiment (Raven-Oncall) nodes. Research conclusions feed into code; code artifacts feed into experiment. The DAG makes the collaboration observable — task launch status, dependency satisfaction, handoff artifacts, and continuation rationale are all explicit.
Memory Connecting Tasks and Experience
After a task finishes, downstream members receive not just a completion signal but also conclusions, files, citations, and context. For backends supporting local file reads, Raven attaches memory record entries from dependency nodes; for session-resumable backends, it reuses continuous context. Cross-session accumulation is handled by EverOS, which stores user context, agent experience, and world knowledge for future reference. Experience must further influence execution strategy — exposed problems, repeated user requirements, and successful practices are translated into Harness adjustments via the modular Harness design.
Harness Self-Evolution: Making One Task the Next Task's Starting Point
The Curator (experimental in V0.2.0) implements the RSI loop at the Harness level. It operates in phases: understand (read current Harness, task, traces, feedback), select (choose which policy surface to modify), design (plan the change), implement (generate code/config/prompt), and fix (repair validation failures). Four policy surfaces — Memory, Planning, Capability, Action — each expose a public protocol; Curator emits Python classes implementing these protocols, bound via an adapter layer to Raven's existing callbacks. Prompts, skill packs, and playbooks are written to the agent's home directory; new tools and gates register as plugins. Every change is inspectable, rollbackable, and auditable.
A simulated travel-planner case (experimental/simulation/cases/s0925c) demonstrates the loop: a fictional agency owner uploads SOPs and role-plays customers. Curator initially generates process and check code (e.g., count question marks per message, reject if >2). Over three feedback rounds, the owner's critiques are codified into Harness adjustments; by round 3 all 11 red-line criteria pass. The repository includes full cultivation logs, per-round code diffs, before/after plans, and reproduction commands.
Caveat: This is a single simulated run with model-played personas; Curator remains experimental. Achieved: multi-round closed-loop Harness rewriting with validation and rollback in one scenario. Not yet achieved: cross-scenario statistical validation or integration with slow model-parameter updates.
From Orchestration to Professional Execution: Capability Evaluation
Raven evaluates both orchestration quality and per-member professional competence across multiple public and internal benchmarks.
Orchestration Evaluation
Metrics: Node F1, Edge F1, Partial Order Accuracy, Exact Match Rate. On Qwen3.8-27B: 0.923, 0.812, 0.950, 0.711. On DS-V4-Flash-0731: 0.963, 0.897, 0.918, 0.867. Both exceed Hermes and Claude Code under the same base models.
Research (Raven-Research)
DeepResearch Mixed (BrowseComp, FRAMES, HLE, xBench-DeepSearch): Qwen3.6-35B 56.3% vs DeepSeek-Harness 49.7%; Qwen3.5-397B 59.3% vs 56.0%; DS-V4-Flash 76.5% vs 68.9%. DS-V4-Flash avg query: 2.38M input tokens, 41.6k output tokens, $0.0242 cost.
Coding (Raven-Code)
SWE-bench Pro (Qwen3.8-27B, offline gen/online score): 54.4% solve rate vs Claude Code 52.4%. SWE-bench Verified (DS-V4-Flash, online): 91.0% vs DeepSeek-Harness 90.4% and OpenCode 90.2%. SWE-Refactor (DS-V4-Flash): avg score 16.5 vs Claude Code 7.0; (GPT5.6-Luna Max): 13.5 vs Codex 10.5. WorkBuddy-Code (DS-V4-Flash): full-set Reward 78.1%. DataAgentBench 2026-08-24 Live (Opus-5): Pass@1 0.8762 (rank 1) vs Permute EQ 0.8713.
Design (Raven-Design)
Slide generation and visual design benchmarks show Raven-Design outperforming Claude Code on the same base model across multiple dimensions (see Table 1 in source).
Continuous Execution (Raven-Oncall)
AI4AI: Nanochat 50M pre-training on 2×A800 80GB. 7 rounds, 172 training runs, zero crashes. Val_bpb relative reduction 5.8% under fixed 20-min single-GPU budget per run. Deliverables include results, visualizations, posters, slides, and project site. AI4S internal benchmark (Opus-5): task success rate 82.35% vs Claude Code 64.71% (Table 3).
Real-World Application Cases
Raven RSI (AI improving AI R&D): Autonomous nanochat pre-training optimization — 7 rounds, 172 runs, 5.8% val_bpb drop, full artifact delivery.
Raven's own release materials: Browser physics mini-game, 16-page product deck, bilingual posters, README — end-to-end multi-capability production.
Godot game asset pipeline: 8-node DAG covering reference research, art spec, parallel weapon/boss asset creation, code integration, design review.
THRESHOLD FPS game: Human-provided PRD; Raven ran ~4 days, 42 planning/development/validation cycles, Godot 4, delivered playable game, posters, slides, site.
Cultivating Your Digital Partner with Playbooks
Users define member roles, node tasks, dependencies, and input arrangements in playbook.md, validated via raven playbook validate. The same Raven-Research backend can assume different research roles (market analyst vs technical researcher) via role-specific task instructions. User knowledge, SOPs, habits, and judgment standards flow into member work styles. EverOS handles cross-session context; SkillForge retrieves skills from local library, EverOS memory, and SkillHub catalog; Proactivity combines event monitoring with scheduled execution; WebUI provides unified task launch, progress tracking, and delivery inspection.
Capabilities for Agent Developers
Upper-layer orchestrator for existing agents (Claude Code, Codex via ACP/CLI/OpenAI API).
Reusable DAG task graphs and Playbooks with parallel branches, failure skip, and validation.
Four ready-to-use professional agents (Research, Code, Design, Oncall) — standalone or in task graphs.
AI-modifiable Harness interfaces: Memory, Planning, Capability, Action policies exposed via home files, config, hooks, plugins, and public protocols.
Reference Harness self-improvement implementation (experimental Curator with phased generation, validation, rollback; s0925c case with full logs).
Cross-session memory and skills via EverOS and SkillForge.
Observable, replayable execution: tracing enabled by default; failed traces exportable for deterministic regression tests.
Raven V0.2.0 is fully open-source under Apache-2.0. GitHub: https://github.com/EverMind-AI/Raven. Cloud version coming to EverMe in October.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
