Harness Engineering for Stable AI Agents: Context, Scheduling & Production Lessons

This article details Harness Engineering practices for building stable AI Agents, covering when to apply Harness, constraining action spaces with skills/MCP, managing context via data precomputation and cost control, implementing instruction sets and event-driven scheduling, persisting memory with version control, and demonstrating measurable improvements in tool-call efficiency and token costs through cache optimization.

Huolala Safety Emergency Response Center
Huolala Safety Emergency Response Center
Huolala Safety Emergency Response Center
Harness Engineering for Stable AI Agents: Context, Scheduling & Production Lessons

What Is Harness Engineering?

OpenAI introduced the concept of Harness Engineering in February 2024. "Harness" translates to horse tack — reins, saddle, protective gear — tools that let the "horse" (the model) run better and more stably. The core demand: make the model "hit exactly where aimed."

Harness Engineering is the deterministic layer wrapped around the model to guarantee normal, stable, efficient operation: constraining action boundaries, scheduling tasks, managing context, persisting state and memory.

When Do You Need Harness?

The author outlines three scenarios where Harness is not necessary:

Scenario 1: Short-term, ad-hoc, low-volume single-turn Q&A or retrieval → model + system prompts + RAG suffices; Harness adds complexity.

Scenario 2: Ample budget to use the strongest model directly → Harness may lower the ceiling.

Scenario 3: Fully AI-native architecture where core value sits on the model's reasoning/generation/decision → small teams can match large-company output; current reality is building assistive plugins around models.

Agent = Model + Harness. Model handles smart; Harness handles stable.

Why Harness Engineering? A Concrete Use Case

The author's team wanted the model to:

Detect malicious attacks: Gateway raw logs reach terabytes per day; even after cleaning, near-terabyte scale. Feeding this directly to the LLM explodes context, cost, and latency.

Assess risky requests: Without MCP tools the model has no "eyes"; with tools, where should its attention focus? Must we remind it every step?

Manage team memory: What deserves long-term memory? How to migrate memory when extending the Agent platform?

How to Do Harness Engineering: Four Pillars

1. Constrain Action Space (Tools / Skills / Permissions / Retry & Circuit Breaker)

Agents run in sandboxes for data security. First step: give the Agent "vision" and "reach" to real data. For small data, manual copy-paste works; for repetitive tasks, build stable, reusable skills and MCP tools . Fixed tools improve cache hit rates and lower cost.

Skills are not "the more the better." Community feedback: installing 10,000+ skills degraded Agent performance. Add skills iteratively, problem-driven. Examples: idp-mcp for Hive queries, ES query skills for real-time traffic, threat-intel skills for IP reputation. These atomic capabilities give the model ready-to-use tools, avoiding ad-hoc reasoning each time, and are reusable across Agents.

2. Manage I/O & Context (Layered Dimensionality Reduction, Cost Control)

Data Precomputation

Don't feed low-information-density data to the model; let the model do high-value reasoning and attribution.

The warehouse has three layers:

ODS (raw logs) — terabyte scale

DWD (cleaned, feature-extracted) — terabyte scale

DWS (statistical dimensions) — hundred-megabyte scale

ODS and DWD are cleaned and reduced outside the Agent. DWS daily data is only ~100 MB; partitioned by hour it becomes a few MB — trivial for the LLM. DWS already contains aggregated dimension and fact metrics. DWD detail data still serves for fine-grained analysis when needed.

Cost Control

Reduce context pressure. After giving the Agent atomic tools, we command it to handle complex tasks. We want more data and longer runs, but attention drift lowers efficiency and tokens aren't infinite. The model must trade off attention.

Passive limits (external):

Tool layer: MCP tools return limited rows, session connection timeouts

Model layer: thinking rounds, thinking duration

Cost layer: token plan

Active interventions (internal):

Control queries: skill docs specify time range, query scope, depth; prompts enforce

Hierarchical analysis: high priority → full analysis; medium → Top-N; low → sampling

Sampling analysis: low priority sampled checks

Covering 1,000+ APIs, the model cannot scan all traffic. Methodology: use cheap resources to pre-compute risk scores at DWD layer, then hierarchical analysis — high risk + ES full query; medium/low risk → layered analysis.

3. Scheduling & Triggering ("When to Work" Handed to Deterministic Systems)

Instruction Set

Simplest approach: inject prompts at system level (e.g., CLAUDE.md, ./prompts). Clear definitions maximize adherence; frequent tweaks destabilize output. One-time solid definition yields highest return.

duties.md example:

Work mode: receive question → think if tool needed → formulate analysis plan → execute tool calls → structured answer

Presentation: structured, agreed format

Principles: which tools, solution comparison, info sources

Prohibitions

rules.md example:

Prevent intermediate files scattered in workspace; keep workspace clean

Stay silent unless @-mentioned; hallucinations hard to eradicate even via memory — declare boundary constraints in rules.md

Task Scheduling

Agent should not decide "when to work"; scheduling belongs to external deterministic systems. Analysis, reasoning, attribution → model.

Typical approach: cron jobs. But upstream task completion times are uncertain → duplicate runs, missed runs, timeouts, low success rate. Tried A2A calls (early unsupported) and scheduled data probes — partial fix.

Switched to event-driven triggering: upstream completion fires event → triggers model execution. Model no longer guesses upstream readiness.

4. State & Memory Persistence (Memory, Git Versioning)

Memory Management

Humans act as reviewers during interaction, contributing business judgments, preferences, workflows — private-domain info. Without this pre-accumulated memory, the "break-in period" is long. Advice: mind memory management from day one as a long-term asset.

Version Control

Use git to manage Agent base files: system prompts, memory files, startup scripts. Ensures traceability across iterations, team sharing, sandbox rebuilds without loss.

Operational Metrics Comparison (Same Model: qwen3.7-plus, Same Period)

Run metrics comparison chart
Run metrics comparison chart
Tool call count comparison chart
Tool call count comparison chart

Key changes driving the difference:

Scheduling mode: shifted to event-pipeline driven, no cron dependency. Agent no longer manages/modifies its own scheduling state → fewer tool calls.

Retry circuit breaker: new rule "max 2 retries per operation, don't repeatedly try different param formats." Previously token explosion and long runs came from model spending time/attention on tool-call failures.

Query method: changed repeated queries to exporting one data snapshot to workspace, then local aggregation — eliminated repeated Hive reads/writes and error retries.

Logic solidified into scripts: script-first; high-consistency tasks frozen as stable scripts/commands, preventing 口径不一致 (inconsistent calibers) and saving on-the-fly LLM reasoning time.

Second run: Agent Loop significantly shortened, stability improved, execution efficiency markedly better. Overall token consumption increased, but most increment from cache reads and cache creation . Cache-hit price is 1/10 of normal input price — small cost for large efficiency gain.

Task completion rate evaluation chart
Task completion rate evaluation chart

Fundamentally, back to the opening principle: let the model focus on solving the problem, not figuring out the problem.

Outstanding Issues & Outlook

Beyond keeping the model running stably, if an Agent only outputs without human feedback it may self-reinforce into a "dead end." Next focus: collect genuine human feedback/corrections, let Agent summarize to avoid repeating mistakes.

Reference reading:

Harness Engineering for Self-Improvement

https://openai.com/index/harness-engineering/

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Memory ManagementAI agentsVersion ControlContext ManagementProduction AIToken OptimizationHarness EngineeringEvent-Driven Scheduling
Huolala Safety Emergency Response Center
Written by

Huolala Safety Emergency Response Center

Official public account of the Huolala Safety Emergency Response Center (LLSRC)

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.