Agent Harness: Writing a Resume That Shows Depth in LLM Agent Projects
This article breaks down the Agent Harness project—a cross-border e-commerce AI employee that validates new products in their first week—to demonstrate how to write a resume for Agent/LLM roles that highlights business value, system design, dynamic tool selection, context management, pluggable tools, stable engineering, and evaluation with concrete evidence.
Cross-Border Store Operations AI Employee: Agent Harness Design
This article dissects a cross-border AI employee that handles new-product first-week validation end-to-end, exposing five core designs: Agent Loop, Context, pluggable Tools, stable engineering, and system evaluation. The system operates under a $300 ad budget, respects gross-margin, inventory, and permission constraints, and delivers per-product continue/adjust/stop decisions with evidence.
Business Story: First-Week Validation Loop
A cross-border team has candidate products but lacks people to continuously monitor them. The AI employee selects up to 3 candidates, completes 7-day validation in a specified market, and delivers investment decisions. Flow: product selection → content listing → test ad sampling → diagnosis & adjustment → inventory handling → retrospective handover.
Business value must hold: Selection, creatives, ads, logistics, and inventory are scattered across systems. Connecting query, judgment, execution, and retrospective makes operational conclusions continuously trackable.
Decisions must have real consequences: When conversion is weak, judge whether it's creative, price, or delivery; when inventory is tight, limit scaling. Execute within authorization; price changes re-verify gross margin; purchasing actions go through approval.
Delivery must influence the next step: Per-product reconciliation of spend and adjustment results, output investment decisions. If end-of-period evidence is insufficient, deliver follow-up plan; extension requires authorization; if no viable product, stop.
Scenario-Driven Tool Selection
Tool selection is driven by the operational problem. Same goal "improve conversion", different causes → different actions:
Listed but no impressions: Check salable status, keyword coverage, ad status. If not yet effective, follow up review; if coverage insufficient, adjust keywords or create test ads.
Enough impressions, low CTR: Pull audience, placement, and creative reports; run main-image or title experiments on suspected issues; reuse qualified creatives directly.
Clicks but weak conversion: Check price, delivery, inventory, and page. If evidence points to high shipping cost, call logistics quoting tool to compare options, recalculate margin and timeliness before adjusting; if page content missing, supplement it.
Inventory near lower bound: Check available stock and inbound ETA, limit scaling or pause related ads, create replenishment request; approval result drives next actions.
Insufficient samples: Verify data window and attribution delay, register scheduled wake-up to continue observing. The correct action may be to wait.
Resume line: Dynamically organized query and operation tools based on impressions, conversion, inventory, and execution anomalies; supported evidence supplementation, creative experiments, placement adjustments, and wait paths; recorded tool selection rationale and execution results.
Agent Loop: Goal-and-State-Driven Decision Loop
The Agent Loop simultaneously manages role goal, current plan, and latest feedback, chaining multiple decisions into one job.
Replan around goal: Persist product, campaign, inventory, budget, and hypotheses to verify. Decompose dependent subtasks; estimate first then price, only after launch can test ads. When new data appears, locally revise plan, reuse existing artifacts.
Dynamic action selection: Each round first identifies current gap, then selects query, execution, or wait from legal capabilities. Model outputs tool, parameters, and expected checks; code verifies permissions, preconditions, version, and budget.
Distinguish continue, wait, and end: Reviews, data windows, approvals all save wake-up conditions. Stop idle loops when no new evidence yet repeated actions; downgrade or handover when budget exhausted. At period end, an independent verifier reconciles delivery.
Goal & State → Supplement Evidence → Select Tool → Execute Verify → Update Plan
Context: Layered Assembly Around Current Decision
Stuffing seven days of reports, creatives, and call logs into the prompt drowns critical constraints. Context must be assembled around the current decision.
Layered organization: Stable layer holds role rules and tool protocols; task layer holds operational goals, budget, authorization; working layer holds current product, campaign, and hypotheses to verify; retrieve historical experiments and external references on demand.
Preserve evidence and caliber: Facts enter structured state; reports label currency, timezone, observation window, and attribution caliber. Long results use summary plus original reference; separate speculation from confirmed facts to avoid writing correlation as causation.
On-demand load and update: Filter by store, product, campaign, stage, and version first, then fetch relevant evidence. Reserve output tokens; when price or inventory changes, re-verify plan and authorization; expired summaries become invalid.
Verification: Compare full history vs. on-demand assembly, simultaneously checking tokens, constraint omissions, stale fact misuse, and task completion rate.
Tool Pluggability: Capability Registry + Business Contract + Platform Adapter
Agent Loop (selects per current task) → assets.render (unified business contract) → Service A or Service B
Replace same-type implementation: Swap image service A for B; Agent still calls assets.render. Adapter converts parameters and results; contract tests verify semantics, format, and error states—identical interfaces are not enough.
Integrate new capability: Add logistics quoting capability; register description, schema, market, permissions, side effects, and cost; complete preconditions and evaluation before entering candidate set.
Filter then select: Runtime filters candidate tools by permissions, stage, and availability; model chooses capability and parameters based on current gap. Identity injected by runtime; returns distinguish completion, processing, rejection, and unknown result.
Boundary selection: After confirmed failure, switch to backup implementation; for write operations with unknown result, verify first.
Stable Engineering: Cross-Day, Cross-System Recoverability
Ad creation may have succeeded but response lost; budget adjustment may face worker restart. The system must know which actions have occurred and which remain uncertain.
Persist and wait for events: Save plan, campaign state, artifact versions, and execution intent. Checkpoint each step; release worker while waiting for review, report, or approval; resume on timeout or event receipt.
Control remote side effects: Timeout enters UNKNOWN; reconcile via platform-supported idempotency keys, request queries, or campaign IDs. Leases and version checks constrain local scheduling; in-flight writes still need remote confirmation—cannot blindly recreate ads.
Manage shared budget: Concurrent products uniformly record consumed, reserved, and pending-confirmation costs; hard budgets use platform-enforced total caps—local reservation cannot replace platform constraints. Traces link goals, tool versions, external objects, and state transitions.
Ad creation timeout handling: Request timeout only means response not received—cannot conclude ad was not created. System retains UNKNOWN, then queries platform state: if already succeeded, reattach to original task; if confirmed not executed and retry conditions met, controlled retry; if still unclear, continue reconciliation or handover. All three branches must synchronously update task, remote object, and budget records to avoid duplicate ad creation or duplicate budget occupation.
System Evaluation: Decisions, Tool Replacement, and Role Delivery
Role tasks must simultaneously verify decisions and execution. Reasonable stop, wait-for-approval, and successful execution all need corresponding acceptance criteria.
Evaluate choice rationality: Cover insufficient impressions, weak conversion, tight inventory, and data delay. Accept allowed actions, evidence, and external states; do not enforce unique call order; observe stale facts, invalid calls, and unauthorized actions.
Evaluate pluggability and recovery: Swap adapter for same capability, rerun contract and task cases; add new capability, check if discovered under appropriate conditions. Inject throttling, restarts, duplicate events, and lost responses after writes; verify task resumption.
Fixed caliber for comparison: Fix initial state, model, and budget; use action-responsive simulated environment for ablation; keep independent test tasks. Report role task pass rate, error side effects, extra takeovers, and total cost per successful task.
Verification: Historical replay only checks decisions. Post-action operational changes must be observed in stateful environment—simulated scorecards; sales and ROAS gains need separate validation in real experiments.
Optimization from Bottlenecks in Role Execution
Optimization starts by locating bottlenecks: repeated decisions, insufficient information, unsuitable tools, or expensive execution. Traces and controlled experiments decide where to improve.
Tool retrieval and ranking: When capabilities grow, recall candidates by task gap, then consider market, availability, and cost; test key tool recall rate and invalid call rate.
Model routing and parallelism: Simple extraction uses small model; complex diagnosis upgrades; independent products run in parallel, same campaign defines single writer, shared budget uses unified reservation.
Role memory and SOP: Distill confirmed store preferences, effective experiments, and operation steps; record source, applicability conditions, and version—cannot generalize one success to all products.
Feedback to training: From failed trajectories distinguish knowledge gaps from decision errors. Persistent behavioral issues then consider SFT/RL; use independent verifier to constrain reward; evaluate generalization and side effects.
Building the Role from Scratch
Start with one market, one store, few products, preserving the full loop from decision to execution, observation, adjustment, and retrospective.
Build runnable environment first: Define product, ad, inventory, and report states; integrate tool adapters. No store permissions? Use stateful simulation platform and virtual clock to advance seven days.
Then build basic role loop: Run through selection estimation, content listing, authorized test ads, data reading, adjustment, and retrospective. Clarify which are code calculations and which need model judgment.
Proactively create branches: Prepare low-click, weak-conversion, insufficient-inventory, tool-offline data; observe if different tools are selected; replace image adapter, then register logistics quoting capability to verify extension.
Finally add recovery and evaluation: Introduce budget contention, lost responses after writes, and restarts; reserve non-tuned tasks for controlled comparison; then add real-authorized environment integration tests.
Project evidence for interviews: Two different execution traces for same goal, one tool replacement and new capability integration, one failure recovery, and one fixed-caliber evaluation report.
Final Resume Structure
Cross-Border Store Operations AI Employee — Agent Harness · [Start–End Date]
Addressed high new-product trial cost and cross-system monitoring needs for [platform/market]; took on new-product first-week validation, completing selection, listing, test ads, diagnosis, and adjustment within fixed budget; delivered per-product investment decisions and execution logs. Integrated with [real authorized store / sandbox / simulation platform]; personally responsible for [explicit scope].
Agent Loop: Designed goal-and-operational-state-driven decision loop; organized plans via task dependencies; selected query, execution, or wait based on impression, conversion, and inventory feedback; supported local replanning and cross-day task progression.
Context Optimization: Assembled rules, operational state, and evidence by store, campaign, and stage; unified report currency and observation window; controlled history growth with summaries and references; maintained fact versions and invalidation mechanisms.
Tool Pluggability: Designed capability registry, standard contract, and platform adapter layer; supported same-type tool replacement and new capability registration; filtered candidate tools by permissions, state, and availability; Agent dynamically selects and validates execution results via contract.
Stable Engineering: Persisted plans, execution intents, and artifact versions; implemented event wake-up, checkpoint recovery, and shared budget management; for write operations like ads, retained unknown state on timeout and reconciled remote results; traced decisions, tool versions, and external changes.
System Evaluation: Built [N] role tasks and [M] anomaly types covering operational branches, adapter replacement, new capability integration, and failure recovery; fixed initial state, model, and budget; conducted controlled and ablation tests in stateful environment; independently verified delivery and side effects.
Measured Results: Role task pass rate [passed/total], baseline change [percentage points]; unit successful task cost [A→B]; extra takeover rate [measured value]; error side effects [count/write operations]. Simulation and real integration reported separately.
Project Artifacts: [Code & Demo] | [Branch Decision Traces] | [Tool Replacement & Integration Logs] | [Operational Retrospective & Evaluation Reports]
Algorithm roles foreground planning updates, tool selection, and context; engineering roles foreground pluggable contracts, cross-day scheduling, and failure recovery. Pick the 3–5 strongest evidence-backed points; delete unimplemented capabilities; fill numbers with measured results.
A project worth an interview must simultaneously show: Business worth doing, decisions have difficulty, system can deliver, results have evidence.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Wu Shixiong's Large Model Academy
We continuously share large‑model know‑how, helping you master core skills—LLM, RAG, fine‑tuning, deployment—from zero to job offer, tailored for career‑switchers, autumn recruiters, and those seeking stable large‑model positions.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
