How an AI Agent Cuts Error Code Triage from Hours to Minutes: Three Practices
This article details how an AI Agent on an orchestration platform automates error code root cause analysis by integrating knowledge bases, code graphs, observability data, and code hosting platforms, achieving 88% consistency with human annotations and zero hard conflicts across 50 test cases through a five-step investigation workflow, dual-path JSON parsing, and regression-tested prompt optimization.
Introduction
SRE teams face hundreds of deduplicated error code alerts daily. Manual investigation of a single error code takes 3–8 hours, with most time spent on cross-system retrieval and context stitching. Even after escalation, developers often re-investigate. To solve this, the authors built an error code governance Agent on an intelligent agent orchestration platform that mounts four capabilities — knowledge base (kb-retriever), code relationship graph, observability platform, and code hosting platform — into a single Agent. The Agent autonomously orchestrates the investigation flow and produces root cause analysis with fix suggestions in minutes. In a 50-case comparison, the optimized action_type consistency with human annotation rose from 66% to 88%, and hard conflict rate dropped from 20% to 0%.
From Alert Forwarding to Fix Suggestion Review
The legacy governance chain was: alert → SRE queries docs/code/logs → escalate to developer → developer re-investigates → fix. The core difficulty is not alert volume but scattered evidence: error codes are opaque numbers, repository mapping is missing, static code doesn't reflect runtime, and SREs can only forward symptoms. The goal is to let the Agent complete root cause analysis and repair suggestions, with SREs reviewing before push: high confidence → "recommend adopt", medium → "reference, needs confirmation", low → archive only, P0/P1 exceptions pushed directly. Developers receive conclusions with evidence and code locations instead of raw alert counts.
Upgrade SREs from "alert porters" to "fix suggestion reviewers".
Why a Single Agent with Four Mounted Capabilities
2.1 Selection: Preserving Full Context for Progressive Exploration
The team compared two architectures:
Multi-Agent collaboration : separate agents for knowledge, code, runtime, with a summarizing agent. Pros: clear responsibilities, suits independent parallel sub-tasks. Cons: communication and orchestration complexity, context loss during handoffs.
Single Agent with all capabilities : one Agent autonomously calls all four tools. Pros: complete context, enables cross-verification. Cons: complex prompt, requires explicit capability boundaries.
Error code investigation matches the second pattern: after finding a handler, trace upstream; when logs show rate limiting, verify error code definition. They chose a single Agent and codified the full investigation flow, tool usage rules, and output constraints into one Skill to control prompt complexity. Splitting into multiple Skills was tried but abandoned because switching relied on Agent initiative and common rules needed duplicate maintenance. Dynamic sub-task orchestration added state management overhead and was not the focus.
2.2 Four Capabilities and Their Evidence Roles
Knowledge base (kb-retriever) : provides repo mapping, business architecture, API docs, historical experience. Constraint: repo_name must come from knowledge base; cannot guess repo from service_name.
Code relationship graph : locates handler, fetches source, traces upstream/downstream calls and error code definitions. Constraint: repo alias only; no enumerating all repos or passing unsupported params. Retrieval path fixed: business repo query → on miss use @<groupName> multi-repo mode → code hosting platform search fallback. Graph call budget tuned from 5 to 15 (5 truncates call chains, 30 risks timeout).
Observability platform : supplies error logs, traces, frontend insights. Constraint: must query when both today_count and today_ret_pct are non-zero; time window uses error_context.date, not default 24h. Query failures skip without retry.
Code hosting platform : line-level blame, commit details, code search fallback. Constraint: only after pinpointing file_path:line; optional, failures skipped and excluded from confidence scoring.
2.3 Orchestration Platform Workflow vs Agent Division
The orchestration platform workflow handles input pre-processing, conditional branching, Agent invocation, result parsing, and output assembly. The Agent performs autonomous investigation and outputs JSON. The Skill carries the investigation process and rules; business side provides knowledge, code, and runtime data.
Workflow is the "skeleton", Agent is the "brain", four capabilities are the "senses".
How Evidence Forms Trustworthy Investigation Conclusions
3.1 Five-Step Investigation, Not Mechanical Serial
Upon receiving error_context, the Agent follows five steps, dynamically choosing query paths based on intermediate results:
Complete error code semantics : if input gcode_meaning empty, verify via at least two sources (docs, code definition, runtime logs); if still unknown, fill ErrUnknown_<retcode>.
Establish business context : retrieve frontend/backend repo mapping, business architecture, API docs, historical experience to seed code localization.
Locate code and call chain : find handler, pull source, trace upstream/downstream and error code definitions, obeying retrieval path and call budget.
Collect runtime evidence : query logs and tag distribution, retrieve traces and full chains, supplement frontend insights when needed.
Supplement recent changes : once a specific code line is located, query blame and commit on demand to populate recent_change.
Each evidence item cites its source (e.g., "knowledge base hit", "code graph context", "observability log", "observability trace", "code hosting blame") for human audit.
3.2 Cross-Validation: Static Code ≠ Runtime Behavior
Static code may show A calls B, but only traces distinguish whether A serially calls B multiple times causing cumulative timeout, or B itself errors after request arrival — two scenarios with completely different ownership and fix actions. Frontend/backend also cross-validate: backend errors check frontend call patterns and error handling; frontend errors verify backend response structure. If frontend source unavailable, explicitly state so rather than infer from backend code. When entry functions delegate error handling, continue tracing instead of stopping at entry.
3.3 Confidence Model
Four evidence sources each score 1 point: knowledge base, code localization, actual source code, runtime evidence. 3–4 points = high, 2 = medium, 0–1 = low. If error code semantics cannot be completed, confidence caps at medium. Code hosting platform changes do not score: knowing who changed code ≠ root cause. Observability logs and traces verify runtime and participate in scoring. Confidence tiers drive the human-in-the-loop push strategy from section one.
11 Nodes for Stable Output
4.1 Fast Path Parsing, Slow Path Fallback
The workflow contains 11 nodes (including Start/End) with conditional branches handling invalid input, normal parsing, and parse failure — not all nodes execute serially.
Start : receives 15 input fields; gcode_meaning optional.
N1 : validates core fields, extracts service_name, normalizes retcode.
N2 : judges input validity, routes to N3 (valid) or N2.1 (invalid).
N2.1 : assembles failed response.
N3 : Agent autonomous investigation, outputs JSON text.
N4 : fast path — four-layer fault-tolerant JSON parsing.
N4-true : passes extract_result as result.
N4.1 : routes based on parse_success to pass-through or slow path.
N4.2 : slow path — LLM extracts 9 fields independently.
N4.3 : uses safe_parse to parse 5 object strings and assemble result.
End : merges three incoming edges, outputs result unchanged.
Agent output may contain Markdown code blocks, surrounding text, or missing fields, so a single json.loads is insufficient. Fast path sequentially tries: direct parse, extract code block, find last object boundary, iterate object start points; any success yields output. On fast path failure, slow path uses LLM to repair format and independently extract 9 fields. Non-empty input yielding all defaults triggers up to 3 retries; only empty input or persistent total extraction failure returns full defaults. Key principle: field independence — code_location empty must not clear existing confidence and analysis_summary.
4.2 Output Contract and Parsing Fallback Synergy
A real case failed all four parsing layers because a Markdown pipe character was escaped as \|, triggering JSON Invalid \escape. The error was inside JSON; stripping prefix/suffix couldn't fix it. Therefore the Skill mandates: double quotes, correct newline escaping, raw pipe |, no backtick escaping, code blocks with language tags, and 17 pre-output checks (6 JSON format, 4 Markdown content, 7 field completeness). Output contract reduces format errors; dual-path handles parse failures — each covers one end.
4.3 Three Engineering Details to Check Early
Unify branch output variables : N4 outputs extract_result, End expects result; N4-true must pass through, else valid parse branch may output empty.
Sync new fields in three places : LLM extraction fields, Python entry parameters, workflow node parameter mapping — all three required. On missing required parameter errors, check mapping config first, not just script.
Batch processing must respect platform rate limits : bulk evaluation may hit minute-level quotas; daily peaks may hit daily quotas. Invocation scripts need rate-limit awareness with wait-retry; higher throughput requires platform coordination.
Workflow pitfalls aren't always in code — they can be in node configuration.
Optimize with Real Cases, Not Rule Stacking
Based on 50-case comparison between pre-optimization v3 and post-optimization v4, five rounds of prompt tuning were run. Beyond retrieval path, call budget, and output contract, key changes focused on judgment order and fix suggestion boundaries.
5.1 Exclude First, Then Judge Whitelisting
"Whitelisting" adds error codes to alert exemption list. v3 listed "whitelist conditions" and "non-whitelist scenarios" side by side; Agent matched looser whitelist conditions first. v4 enforces explicit three-step: first check prohibited whitelist scenarios, exclude all, then evaluate whitelist conditions, finally distinguish whitelist vs whitelist_and_refine. Key rules: rate-limit errors prioritized as rate_limit_internal/external; fallback generic codes must refine via fix_code before whitelist_and_refine; timeouts prioritize code and config investigation, only choose handle_cloud_resource when cloud resource itself fails and code/config cannot optimize; expected business interceptions combine business rules, not just code_type.
5.2 Fix Suggestions Need Boundaries
Agent tends to over-suggest: add if-else at every call site, modify non-triggering entry points, extend error code system for low-frequency flakes. Three principles added: modify only actually problematic entry points; prefer reusing existing centralized handling tables or middleware; evaluate user handling divergence and ROI before adding error codes or abstraction layers. In multi-factor scenarios, action_type picks fastest mitigation; other optimizations go into steps, separating short-term fix from long-term refactor.
5.3 Two Representative Cases
Rate limit error 30110013 : v3 → whitelist; v4 → rate_limit_internal. Logs showed rate limit counter triggered; rate-limit check moved before whitelist check, so even code_type=exception cannot skip.
Redis/network timeout : v3 → handle_cloud_resource; v4 → fix_code or fix_config. Trace revealed Redis single call not slow; handler internally serial-called Redis multiple times causing cumulative timeout; prioritize code bottleneck and timeout config investigation.
Such tuning extracts rules from error paths so future similar issues benefit, rather than memorizing per-error-code answers.
Skill engineering isn't writing more prompts — it's comparison-driven, case-driven, rule-explicit.
Results: From Phenomenon Forwarding to Executable Conclusions
6.1 50-Case Version Comparison
v3 = original, v4 = after 5 prompt tuning rounds. Evaluation metrics:
Consistency with human annotation : action_type match rate: v3 66% (33/50) → v4 88% (44/50).
Human intervention rate : action_type mismatch requiring SRE review: v3 34% (17/50) → v4 12% (6/50).
Hard conflict rate : fix suggestion completely contradicts human judgment, unpublishable: v3 20% (10/50) → v4 0% (0/50).
Net improved cases : v3 wrong & v4 correct: 11 cases.
No hard conflicts in this 50-case batch; production still uses evidence-sufficiency grading. The 50 cases served both tuning and regression validation, not generalization proof for new error codes.
6.2 Production Run and Collaboration Shift
Workflow deployed on orchestration platform; past month processes hundreds of deduplicated error code alerts daily, covering backend and frontend H5/mini-program. High confidence ~90%, failure rate ~1% (mostly workflow API rate limits). Governance chain shortened from "hours investigation + days fix" to "minutes investigation + 0.5–2 days fix".
Collaboration changed directly: Agent provides code location, evidence, and fix action; SRE reviews and decides push; developer confirms and executes. Compared to receiving a raw symptom description and re-assembling context, developers continue verification from existing analysis, reducing duplicate investigation and context switching.
Evaluation Loop for Continuous Iteration
Post-launch, every bad case may trigger a new rule. Without regression evaluation, Skill degrades from "increasingly accurate" to "increasingly bloated and slow". Evaluation must cover both conclusion and process.
7.1 Decompose Investigation into Checkpoints
Final result : action_type matches human label.
Process completeness : semantic completion, knowledge base, code localization, runtime query executed per conditions.
Tool reliability : parameters and call budgets correct; explicit degradation on query failure.
Evidence and boundaries : source code pulled, counter-evidence checked, confidence matches available evidence.
Efficiency and stability : latency, workflow failure rate, rate-limit trigger frequency.
Only checking final answer cannot distinguish "correct with evidence" from "lucky guess". Production traces, tool invocations, and human confirmation labels must enter evaluation corpus to avoid irreproducibility after alert/monitoring data expires.
7.2 Three Landed Mechanisms
Version comparison : align v3/v4 by request_id on action_type, generate migration matrix and summary, compare against 50-case, 7-category human labels to see improvements and regressions, not just overall consistency.
Regression case solidification : save key evidence, prohibited error directions, and expected final judgments in Skill workspace; rule iterations must not arbitrarily shift baseline. The 6 cases where "v3 covered v4" (2 empty shells, 2 improper whitelists, 2 judgment divergences) retained for future regression.
Bad case feedback into Skill : abstract concrete failures into generic constraints: cannot lock root cause before querying runtime evidence; cannot whitelist fallback codes mixing multiple fault types; cannot retry tools blindly after empty returns; cannot cherry-pick evidence supporting current conclusion.
Evaluation does not block alert response, but every diagnosis must leave replayable, measurable, extensible artifacts. This gives Skill changes evidence and regression safety.
Conclusion: Reusable Method, Not Entire Rule Set
The transferable core is cross-validation of business semantics, static code, and runtime evidence; structured output via workflow; continuous iteration via bad cases and regression evaluation. Error code semantics completion, action_type enumeration, and confidence scoring need business-specific tuning, not direct copy.
Current work: deepen frontend error handling chain tracing; monitor if 15 graph call budget covers complex scenarios. Next: further refine judgment boundaries; phase two adds multi-error-code batch processing to boost daily governance throughput.
Launch is not the end — it's the starting point for the next iteration cycle.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Tencent Technical Engineering
Official account of Tencent Technology. A platform for publishing and analyzing Tencent's technological innovations and cutting-edge developments.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
