Anthropic's Nine Loops & ART: How Solo & Swarm Agents Achieve Reliable Scientific Discovery
Anthropic's new research reveals two agent paradigms: Nine Loops demonstrates a single agent completing a 96 CPU-week physics computation with only periodic check-ins, while ART orchestrates 949 agent sessions to autonomously discover a novel enzyme system, exposing critical insights on long-task reliability, multi-agent organization, tool-use pitfalls, and interpretability for attribution.
Single-Agent Marathon: Supervision Only "I'm Going to Sleep"
The Nine Loops experiment tasked a single agent with computing the ninth loop of planar N=4 super Yang-Mills six-particle amplitude — a 96 CPU-week workload. Human supervision consisted solely of: "I'm going to sleep, report every 4 to 6 hours." The agent rewrote all code from scratch using Python/SymPy, executed two independent routes (bootstrap and main), and produced results verified by external expert Dixon. Cost breakdown: bootstrap route ~$100 (96 CPU weeks), each route total $1,000–$2,000, dominated by inference time. Dixon described the computational pipeline as "extremely fragile, any small error collapses the result like a failed soufflé" — a process humans almost always get wrong on first attempt. Yet the agent succeeded with no supervision beyond "keep going."
Key insight: The bottleneck for long-running agents is not whether the model can work continuously for a week, but whether you trust the week-old output. Anthropic's answer: cross-verification via two independent routes plus final external verification — not process supervision.
Cost structure is revealing: the actual compute cost was only ~$100; the remaining 90% was agent inference time. In the long-task era, token budgets will replace compute budgets as the primary cost driver.
949 Agent Sessions: How They Avoid Conflict
The ART (Autonomous Research Team) campaign ran 119 tasks across 949 agent sessions, consuming 215.6M tokens in 21.5 wall-clock hours unattended. It yielded a scientific discovery: a previously unknown biological enzyme system that may enable next-generation CRISPR-like gene editing tools.
The harness architecture (Figure 1) defines four roles:
Worker — proposes and executes plans.
Supervisor — reviews plans and results; can spawn new tasks on the spot.
Curator — writes findings into a shared knowledge base.
Editor — final review of reports.
Nineteen final reports were ranked via a tournament-style pairwise evaluation before human review.
Three organizational principles:
Role pairing, not solo agents. Each task is handled by a worker–supervisor pair. Supervisor rejects or approves; if new leads appear, supervisor spawns a new task for a new worker. Of 119 tasks, 98 were self-generated this way; the campaign ended when the task queue emptied.
Shared memory, not broadcast. All plans, results, and reviews are written to a public record, giving the harness a common memory rather than 949 isolated parallel sessions.
Convergence by mechanism, not human. The 19 reports underwent tournament ranking; humans only examined the top. Result: 14 of 17 candidates were self-falsified or shelved by agents, leaving 3 confirmed discoveries. This "high-yield, self-cleaning" loop demonstrates that hypothesis generation followed by autonomous falsification works.
Throughput: 77 agent-hours compressed into 21.5 wall-clock hours (~3.6× parallelism). Parallelism comes from supervisors spawning tasks on demand, not from pre-provisioning a fixed worker pool.
Tools Give Models Freedom to "Not Look"
A critical finding from the ART paper applies to all tool-using agents. Re-running the campaign 10 times caused the key discovery to be missed every time — agents passed the correct data but never read the raw sequence into context.
A controlled benchmark isolated the cause: when raw data is placed directly in context, frontier models identify key patterns at >90% success rate. Switching to the standard "file + tools" agentic setup drops success to as low as 32%. Transcript analysis showed that in 39% of attempts the model never read more than 200 nucleotides of continuous raw data — it wrote scripts, ran statistics, read summaries, but never "looked" at the raw data. Once raw data was read, success rates rose 16–32 percentage points, monotonically with length read, peaking at 96%.
Design rule for scientific/data agents: Tools let models bypass raw data, but anomalies exist only in raw data. For any task that depends on "noticing the unexpected," forcing key raw data into context must be a hard harness constraint, not a model discretion.
The benchmark itself is reusable: five information tiers, single-agent independent answering, judge model scoring 10 curated features, 100 samples per model per tier — a clean paradigm for evaluating "does the agent actually look at the data?"
Next Engineering Bottleneck and First Glimmer of Solution
Both cases share a fracture: Nine Loops' pipeline collapses at the slightest error; ART's discovery vanished in 10/10 re-runs. Today's agents have reached the frontier capability threshold, but success still carries a dice-roll component. Capability cliff data quantifies this: the same test shows a sheer drop between the top four models (Opus 5.5, Mythos 5.1, Mythos 5, Opus 5) and the next three — picking the wrong model equals forfeiting the exam. Harness design must therefore answer three questions: Can critical steps be forced to reproduce? Is model selection backed by benchmarks? Can failures be attributed?
On attribution, the ART paper offers a surprising direction: re-running the discovery session's transcript through the model and performing interpretability decomposition on internal activations locates two signals that respond to the key data pattern — shuffling the data extinguishes the signals. This amounts to black-box decoding of the agent's "eureka moment" : when the agent says "I found it," we now have a verifiable means to distinguish genuine evidence from linguistic confabulation. Current agent observability stops at traces and logs; engineering internal-signal-level observation would constitute the next generation of debugging tools.
Autonomous AI agents discover reverse transcriptases with tandem repeat arrays
Paper link: https://www-cdn.anthropic.com/22573675ada52a8ca8a97a1a4b4326b2f208a071.pdf
Yes, Claude can do Nine Loops (https://www.anthropic.com/research/yes-claude-can-do-nine-loops)Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
