Anthropic's Claude Agent Engineering: Multi-Agent Workflows, Skills & Eval

Anthropic publishes its internal Claude agent engineering practices on claude.dev, covering dynamic multi-agent workflows, hundreds of reusable skills, context engineering principles that cut system prompts by 80%, and evaluation-driven hillclimbing that boosted accuracy to 90.5% at one-fifth the cost.

PaperAgent
PaperAgent
PaperAgent
Anthropic's Claude Agent Engineering: Multi-Agent Workflows, Skills & Eval

Anthropic has launched claude.dev, a developer portal that systematically shares the internal methodology used to evolve Claude from a solo model into a collaborative multi-agent team. The site is organized into four pillars that map directly to the layers of agent engineering: AGENTS, SKILLS, ENGINEERING, and PLAYBOOKS.

AGENTS: Dynamic Multi-Agent Workflows in Claude Code

Instead of pre-written static harnesses, Claude Code now supports dynamic workflows : for each task, Claude composes and directs a custom multi-agent team on the fly. The article identifies three failure modes of single-context agents:

Agentic laziness — e.g., a security review of 50 items stops at 35 and declares "done".

Self-preference bias — when asked to score its own output, Claude is consistently lenient.

Goal drift — after repeated context compression in long tasks, constraints like "don't do X" silently disappear.

The solution rests on three primitives: agent() to spawn a sub-agent, parallel() for fan-out concurrency, and pipeline() for sequential chaining. Each sub-agent receives an independent, clean context window with a single goal, preventing cross-contamination.

A harness for every task: dynamic workflows in Claude Code
A harness for every task: dynamic workflows in Claude Code
Comparison of static harness vs dynamic workflow
Comparison of static harness vs dynamic workflow

Six composable patterns are provided: classifier routing, fan-out–synthesis, adversarial validation (each worker paired with a critic), generate–filter, tournament ranking, and loop-until-convergence. The Bun rewrite from Zig to Rust was accomplished using this orchestration.

Six workflow patterns for combining Claude agents
Six workflow patterns for combining Claude agents

Security is architected in: when handling untrusted input (e.g., public user feedback), a reader agent has read-only access while the actor agent sees only summaries , creating an architectural isolation layer rather than relying on prompt instructions. Combined with /loop, triage teams can run 24/7.

Security architecture: reader agent read-only, actor agent sees summaries
Security architecture: reader agent read-only, actor agent sees summaries

The trade-off is explicit: workflows consume significantly more tokens and suit complex, high-value tasks ; a deep research run may involve 22 agents and 1.1 million tokens. The principle: parallelism and specialization must earn back their coordination overhead.

Token cost trade-off: workflows burn more tokens
Token cost trade-off: workflows burn more tokens

SKILLS: Hundreds of Internal Skills Distilled into Nine Categories

Anthropic actively uses hundreds of skills , which fall into nine categories: library & API reference, product validation, data fetching & analysis, business process automation, scaffolding templates, code quality review, CI/CD, on-call runbooks, and infrastructure operations.

Lessons from building Claude Code: How we use skills
Lessons from building Claude Code: How we use skills
Nine categories of Anthropic internal skills
Nine categories of Anthropic internal skills

Three counter-intuitive insights:

Validation skills deliver the highest ROI — worth spending an engineer's full week to perfect.

Skill descriptions are not human-readable summaries but trigger conditions for the model — they must contain keywords that help the model decide "who handles this" at session start.

Gotchas (pitfall records) are the highest-signal content in any skill and should accumulate continuously with use.

Gotchas section grows continuously with usage
Gotchas section grows continuously with usage

A skill is not a single Markdown file but a folder: SKILL.md acts as a hub pointing to specific files — the entire filesystem becomes progressive-disclosure context engineering .

Skill folder structure: SKILL.md as hub pointing to specific files
Skill folder structure: SKILL.md as hub pointing to specific files

Distribution and governance: small teams commit skills directly to .claude/skills in the repo; at scale, an internal plugin marketplace lets teams install on demand.

ENGINEERING: Deleting 80% of System Prompts Made the Model Stronger

For Claude 5 generation models, Anthropic removed over 80% of Claude Code's system prompts with no measurable drop in coding benchmarks . Many old rules were "casts" for older models (e.g., "never write multi-line comments" because old models got them wrong); new models have judgment, and the casts became constraints. Internal transcripts even showed system prompts, skills, and user requests conflicting.

The new rules of context engineering for Claude 5 generation models
The new rules of context engineering for Claude 5 generation models
Comparison of six old rules vs new rules
Comparison of six old rules vs new rules

Six "old → new" inversions:

Give rules → Give judgment

Give examples → Design good interfaces (examples restrict exploration space)

All upfront → Progressive disclosure

Repeat emphasis → Concise tool descriptions CLAUDE.md as memory → Automatic memory

Simple specs → Rich references (HTML mockups, test suites, rubrics can serve as specs)

Assembled context layers
Assembled context layers

Practical advice: keep CLAUDE.md lightweight, spend tokens on codebase-specific gotchas; bind system prompts to product context (most valuable when building custom harnesses); prefer code references over natural language — an HTML mockup conveys design better than text or a screenshot . Run /doctor in Claude Code to slim down CLAUDE.md and skills.

PLAYBOOKS: Turning "Feels Better" into Measurable Numbers

This section answers "how to prove it works." A good eval requires four properties: task distribution matches production, stronger models score higher, frontier models still have headroom, and low variance across runs .

Automating eval design and hillclimbing with Claude
Automating eval design and hillclimbing with Claude
Four elements of a good eval
Four elements of a good eval

A common pitfall is adversarial sampling : selecting cases because "today's model fails on them" measures the model's failure fingerprint, not task difficulty. Correct approach: humans must articulate why a case is hard; cases come from production traffic, bug reports, and tickets. Scorers are chosen by output space: programmatic verification for closed outputs, LLM-as-judge for open outputs, and the judge model must not be the model under test .

Hillclimbing includes anti-overfitting guardrails: split eval set into train/test; if train improves but test flatlines, declare overfit and roll back the patch ; never paste failing examples verbatim into prompts to avoid reward hacking.

Cost-oriented hillclimbing process
Cost-oriented hillclimbing process

Real results: on an internal support benchmark, starting from Opus 4.8 high-effort (74.4% accuracy, $4.60 per ticket), hillclimbing reduced model size, effort, and prompts to achieve 90.5% accuracy on a held-out set at roughly one-fifth the original cost . The claude-api skill itself was hillclimbed from 66% to ~88% pass rate, uncovering two evaluation bugs where the task spec was wrong.

Many featured videos, recommended
Many featured videos, recommended

Closing the Loop

Anthropic's methodology forms a closed loop: use workflows to team up, skills to codify experience, new context engineering to unshackle the model, and eval to verify every step . The evolution from solo Claude to collaborative agent teams stems not from a single breakthrough but from this compounding engineering discipline.

https://claude.dev/
https://x.com/ClaudeDevs/status/2105391694741119047
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agentsevaluationmulti-agent systemsskillsAnthropicClaude Codecontext engineeringhillclimbing
PaperAgent
Written by

PaperAgent

Daily updates, analyzing cutting-edge AI research papers

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.