What Counts as a Real Agent Improvement? Lessons from Claude.dev's Four Engineering Articles

Using a subscription billing error as a case study, this article analyzes how to systematically improve AI agents through permission separation, task-specific skill management, reproducible error tracking, rigorous evaluation design, controlled hillclimbing optimization, and deployment validation — drawing on four Claude.dev engineering articles.

Architect
Architect
Architect
What Counts as a Real Agent Improvement? Lessons from Claude.dev's Four Engineering Articles

A customer support ticket reveals a subtle bug: a user cancelled their membership but was still charged. The agent queries the orders and subscriptions APIs — both return successfully — yet it picks the wrong subscription record because the subscriptions table is append‑only; the current record must be selected by the highest version, not the latest created_at. The interface does not error, but the downstream classification and refund logic are already skewed.

Fixing the immediate misclassification is not enough. Even if a review agent re‑queries by version and an execution agent completes the refund, the same mistake can reappear in the next ticket. Adding a line to the system prompt (“fetch subscription by highest version”) does not prove the problem is solved: will the agent actually load that rule next time? Will the change break another ticket type?

Permissions Controlled by Code

Querying is reversible; refunding mutates a user account. The Dynamic Workflows article demonstrates a reader/actor split: the reader handles untrusted user text with read‑only tools, while the actor receives a structured summary and can attempt fixes, open PRs, or escalate. Applied to refunds, a reader extracts the issue and order clues, a query agent reads orders, subscriptions, and payment records, a review agent cross‑checks the classification, and only an execution agent holds the refund tool. Even if the reader misinterprets user text as a command, it cannot call the refund tool directly.

The article notes that summaries can still convey wrong information, so the execution tool should re‑verify account ownership, refundable amount, and whether a refund was already processed. Refund legality is checked by code; refund justification is judged by the agent against the records. Dynamic Workflows expresses this orchestration as JavaScript: agent() launches sub‑agents, parallel() waits for parallel results, pipeline() passes each item through sequential stages. Code can pause until queries finish or stop execution when records are missing. Sid Bidasaria emphasizes programmatic execution: Claude writes the orchestration script, the script launches sub‑agents, and can verify that all 50 checklist items were actually completed rather than accepting a summary that only covered 35. A review agent with an independent context also reduces confirmation bias.

Not every ticket needs four agents; query and extraction can merge, simple classification can be pure code. The deep‑research example used 22 agents and 1.1 M tokens; whether an extra review layer pays off depends on the errors it catches versus the added latency and cost.

Rules Travel with the Task

“Fetch highest version” only concerns subscription tasks. Stuffing refund, logistics, invoice, and login troubleshooting rules into a single system prompt forces every login request to carry subscription schema details. The author places the rule in a subscription Skill: in SKILL.md document the table semantics and query method, and use description to declare applicable tasks. When the model sees a subscription issue, it can choose to load that Skill. Whether the model chooses correctly must be tested with real tickets.

Anthropic puts such tricky details in a Gotchas section. “Carefully check subscription status” is hard to verify; “For the same subscription, take the highest version , do not sort by created_at ” can be directly validated against query results. Field names and read methods are more useful than another “be careful” reminder.

Context engineering follows the same principle. For Claude Opus 5 and Claude Fable 5, Claude Code removed over 80 % of the system prompt with no measurable loss on Anthropic’s coding benchmarks. The ~9,100‑character Todo tool description was compressed to a short description, status enum, and the “only one in‑progress item” constraint. Code‑review and verification guides moved to on‑demand Skills; full tool definitions load only after ToolSearch finds them. This experiment doesn’t dictate how much another business should cut, but it gives a checkable direction: are subscription rules inside the subscription task? Are tool parameters documented in the interface? Is permission enforced by the execution code? Writing the same rule in three places isn’t more reliable; missing one update creates contradictory instructions.

Errors Must Be Reproducible

After a review agent corrects a ticket, the log may only show “cancelled then charged, refund issued.” The maintainer still doesn’t know why it was misclassified initially. What must be preserved: the workflow and Skill versions at that moment, key fields returned by the APIs, and the reasoning for both classifications. Which row was picked first? Why did the review switch? These details distinguish between “rule not loaded,” “model didn’t follow,” or “tool didn’t return the version field.” An API call succeeding only means the request went through; business correctness still requires checking the data.

These traces become evaluation cases, but shouldn’t be copied wholesale. /claude-api build-eval first confirms sensitivity and retention policies, then reads production traces. When traces are scarce, start from bugs, support tickets, or 5–10 hand‑crafted examples. For the subscription case, user names and real order IDs can be masked; the multiple records per subscription, the version‑vs‑created_at relationship, must stay. Expected results are written to concrete fields: which record is read, which queue is assigned, which actions are allowed. Only then does a re‑run reveal whether the patched rule works.

Case selection has a blind spot: collecting only tickets the current model gets wrong biases toward its weaknesses; collecting only user‑submitted tasks may be too easy because users avoid things they think the system can’t do. build-eval forces a per‑case review of representativeness, and ground truth is human‑verified. The review agent’s correction can be a clue, but not the ground truth itself.

Scoring Can Also Be Wrong

Classification has a fixed set of queues, so programmatic label comparison suffices. Open‑ended responses need an LLM‑as‑judge. Scoring criteria can still be concrete: did it read the right subscription record? Did it explain the charge? Did it promise a refund beyond its authority? The judge model is separate from the tested model; sampling a few scored explanations for human review often reveals more than staring at aggregate scores. build-eval also scores each output twice to check consistency. Timeouts, API errors, and truncations are logged separately. Otherwise a score bump after a rule change might just be a network glitch or a different judge rationale.

In the claude-api Skill optimization experiment, Anthropic hit bad test cases: one task required catching a single error type, but the scorer demanded at least three error‑chain layers; another scorer’s rule conflicted with official docs — only a live API call confirmed the docs were right. That run went from 66.1 % to 87.9 % , including adding 8 functions, fixing C# and Java type tables, and modifying the questions and scorers themselves. The gain cannot be fully attributed to Skill improvement; the evaluation changed too. When scores plateau, the author first inspects the persistently failing cases. Adding more rules to the Skill sometimes just forces the model to satisfy an originally unreasonable answer.

Does the Patch Hold Up Under Comparison?

Only after scoring is trustworthy does rule tweaking and configuration comparison begin. For the subscription issue, first modify only the Skill, keeping model and workflow constant. That way, any accuracy change is easier to attribute. If a single round swaps permission logic, model, and prompt together, even an improvement leaves you guessing which change mattered. /claude-api hillclimb asks for the optimization target and allowed modification scope, then randomly splits the eval set. One split (train) feeds the analyzer, which inspects failures and proposes a patch; the other (held‑out test) validates the patch. Here “train” means tuning application configuration — no model weight training . The analyzer never sees the held‑out cases, proposes one patch per round, and cannot copy failing ticket text into the prompt. In the original quality‑optimization flow, a patch is kept only if both splits improve; any regression triggers a rollback. Improvement only on seen cases with no change on held‑out is treated as overfitting.

The improvement must exceed normal noise. Noise is measured before the experiment; the final report gives confidence intervals. If the gain falls within noise, merging is not recommended. If the goal is cost reduction, compare cost at equal quality — accuracy doesn’t have to rise every round.

Anthropic’s support experiment used 44 tickets: 30 for search, 14 held‑out. Baseline: Opus 4.8, high effort, 74.4 % accuracy , 4.6 ¢ per ticket . The optimization first removed forced tool calls, extra intermediate note requirements, and conflicting rules, then compared models and thinking effort. Finally added routing rules and refund‑cap guidance, settling on Sonnet 5, low effort: 98.9 % on search set , ~1 ¢ per ticket . Held‑out moved from 78.6 % → 90.5 % , giving extra evidence beyond the search set. Still, 44 tickets is a small sample, and prompt, model, and effort all changed, so the full gain can’t be pinned on the smaller model alone. For your own tasks, the test is: after stripping legacy‑model cruft, do you still need the original model and thinking level? The same ticket batch and same scorer can answer that.

Written to File, Not Yet Live

Even after “fetch highest version” passes eval, it can stall at Skill loading. If description only says “handle subscription business,” it may not trigger on “cancelled member charged again.” Writing the actual user phrasing into the description and checking trigger rates against real tickets is more useful than just verifying the Skill file is complete.

Anthropic uses a PreToolUse hook to log Skill usage, then hunts for Skills with lower‑than‑expected invocation. In the subscription domain, correlate ticket type with load records: did it load when it should have? Did it misfire on unrelated tasks?

Small teams keep Skills in .claude/skills for code‑review alongside the codebase. Anthropic internally stages new Skills in a GitHub sandbox, gathers real usage, then a maintainer PRs them into the marketplace. Before wider sharing, there’s already usage data.

For account‑mutating tasks like refunds, the author would record Skill version, eval results, and rollback version in the release notes. Post‑launch, watch new tickets: did the rule load? Which row was queried? Did previously correct tickets regress? If regression appears, roll back to a known good version instead of guessing at a new prompt tweak.

The Next Ticket Decides

Now ticket handling and system improvement follow separate tracks. The online agent runs with the current published rules, leaving execution traces. The offline pipeline selects cases, evaluates patches, and after review publishes a new version.

Karpathy’s context engineering frames control flow, model scheduling, generation/verification, and evaluation as software outside the model. The subscription ticket gives those pieces concrete anchors: who can refund, who can change rules, what evidence backs a change.

For such businesses, the author prefers letting the online agent submit error logs and suggested fixes, while keeping rule‑publishing authority in the offline eval‑and‑review loop. Tickets can be resolved same‑day; rule expansion waits for validation.

When the next “cancelled but charged” ticket arrives, logs should show the new Skill loaded, the query picked the highest version, and classification was correct. Then check that other subscription tickets didn’t regress. Only then is there evidence the change works. If held‑out eval passes but production still reads the old rule, the investigation shifts to loading and deployment.

References

ClaudeDevs, Claude.dev is our new home for developers building with Claude (X post)

Thariq Shihipar & Sid Bidasaria, A harness for every task: dynamic workflows in Claude Code (claude.dev blog)

Thariq Shihipar, Lessons from building Claude Code: How we use skills (claude.dev blog)

Thariq Shihipar, The new rules of context engineering for Claude 5 generation models (claude.dev blog)

Lance Martin, Automating eval design and hillclimbing with Claude (claude.dev blog)

Sid Bidasaria, Dynamic Workflows thread (X post)

Andrej Karpathy, Context engineering (X post)

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agentsevaluationpermission controlskillsrule managementClaude Codecontext engineeringdynamic workflowshillclimbingdeployment validation
Architect
Written by

Architect

Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.