Map Road Digital Employee: Multi-Agent Platform Automates Road Network Data Analysis 6-10x Faster
This article details a multi-agent platform called Map Road Digital Employee that automates map road network data analysis, reducing case processing time from 20-30 minutes to 3-5 minutes through a four-layer agent architecture, skill-based tooling, state-machine orchestration with human-in-the-loop, and checkpoint persistence, achieving 6-10x efficiency gains while maintaining reliability.
Challenges in Map Road Network Data Analysis
Map road network data production involves many elements, long workflows, and complex application strategies. Teams face three core pain points: high barrier to entry, low analysis efficiency, and scattered reference platforms. A typical user issue analysis requires six key steps across 4-6 different platforms, taking 30 minutes to several hours per case, with hundreds of person-days invested monthly yet low closed-loop efficiency.
Why an Agent-Based Approach
The analysis process is semi-structured: overall steps are fixed but details contain ambiguity. The evolution from LLM to agent systems is traced through four stages, with the state-orchestration stage providing three capabilities that directly address the challenges: semantic understanding and judgment by the model, process control by a state machine, and reusable execution units encapsulating platform query capabilities. This leads to a design where deterministic steps are handled by a rule engine/state machine, uncertain steps by the LLM, and humans intervene only at high-risk nodes.
Multi-Agent Architecture
Four-Layer Hierarchy
The system adopts a four-layer architecture with judgment concentrated in upper layers and execution in lower layers. Reliability is ensured by the state machine and event chain: retry limits with escalation to human, checkpoint persistence for resume, and mandatory human confirmation for high-risk operations.
Six Core Roles
Planner : decomposes tasks into steps
Reactor : judges step outcomes (retry, re-plan, rollback, proceed, abort)
ClarifyAgent : passive clarification when information is insufficient
Summarizer : aggregates conclusions
SkillExecutor : sole executor of real operations via Skills
GeneralAnswer : handles steps without Skill coverage, isolates specification queries
All six roles share a unified BaseAgent base, ensuring consistent lifecycle, hooks, and logging. Scenario-specific SubAgents (case analysis, full-chain retrieval, Kunpeng vertical) are built on the same base.
Skill System and Capability Scheduling
Atomic Skills with Standardized Specs
Map production domain capabilities (e.g., query turn-restriction status, convert tile coordinates, render base map) have deterministic I/O and no complex reasoning, so each is implemented as an independent Skill. Skills are reusable across scenarios, independently verifiable, and modifications don't affect others. Professional knowledge (TTFA encoding bits, tile conversion formulas) is documented in Skill specs, shifting domain expertise from agents to the Skill library.
Skill Spec and Directory-as-Registry
Each Skill exposes a single public interface: a six-section spec covering purpose, invocation, and failure handling. Specs are grouped by business domain in a directory tree; at startup the directory is scanned into an in-memory registry. Adding a Skill only requires dropping a new spec file — no central registry updates, eliminating config drift.
Capability Manifest and Closed-Loop Execution
The scanned registry generates a unified capability manifest (name, description, invocation spec, risk level) injected into Planner, Reactor, and ClarifyAgent, ensuring consistent understanding. Execution follows a two-step dispatch. The task advances in a closed loop: any step's evidence may need re-fetching, so state-controllable rollback is a prerequisite for automation. Conclusions must reference the evidence layer and indicate the next production platform.
Judgment-Execution Boundaries
Four boundaries separate judgment (agent) from execution (Skill): deterministic actions sink to Skills, judgments float to agents, and data level (SD vs LD network) is confirmed first.
Orchestration and Human-in-the-Loop (HITL)
Finite State Machine Topology
The Supervisor orchestrates a session as a 7-node, 5-condition-edge FSM: planner → confirm_gate → execute_step ↔ rollback → clarify / human_interrupt → summarize. All edge decisions derive from Reactor's conclusion, keeping state explicit and judgment centralized.
State Management
State is limited to two groups: four transient fields (retry position, pending rollback, pending output, pending user confirmation) and core scheduling fields (current step index, result list, remaining rollbacks, current step input, history, available tools). All state must be serializable; checkpoint resume is a design constraint, not an add-on.
Three-Layer Guards
To prevent runaway: per-step retry limit, per-session rollback limit, and global recursion depth limit. When the system cannot progress, it halts and hands off to a human — reliability over completion rate.
Two HITL Modes
Confirm (active) : before risky execution, system asks for approval; deterministic steps skip this.
Clarify (passive) : when stuck, system asks user to supply missing evidence.
In map scenarios, turn-restriction modifications default to confirm; missing version history triggers clarify.
Suspend and Resume
Sessions suspended for HITL or failure support resume via checkpoint snapshots. A data modification confirmation may span hours; on return the system continues from the confirmation point without re-executing prior steps.
Engineering Implementation
Three-Repository Structure
Physical layout:
engine</core> (orchestration, state machine, BaseAgent), <code>skill(Skill implementations and specs), service (Go backend for persistence, callbacks, compensation). They communicate via a unified HTTP event protocol, independently deployable.
LLM Factory and Model Abstraction
Multi-vendor access (Wenxin, GLM, GPT, Claude) is solved by a registry factory keyed by model name. All code depends only on a stable BaseLLM interface (single-turn completion, multi-turn chat, multimodal). Auth, timeouts, context differences are encapsulated per implementation; switching vendors is a config change.
BaseAgent Template Method
Six roles inherit BaseAgent, differing only in run(). The base uses template method pattern with three hook slots ( on_start, on_end, on_error) for observability; on_error reports but does not swallow exceptions, preserving Reactor's ability to judge success.
Constants and Global Config
Values needed in multiple places (engine routing, protocol mapping, frontend display) have a single definition source to prevent drift.
Observer Pattern for Hooks and Events
Observability via AgentHook (three methods, separate package, no circular deps). Built-in AutoLogHook (execution summaries) and StateHook (translates to external event stream, per-run construction avoids cross-talk). One internal event can map to 0-N external events (e.g., terminal splits into message and run_finished so frontend never waits indefinitely). Event flow crosses process boundaries via Go service callbacks acting as remote hooks.
Runtime Mechanisms
Graph State and Serialization
Graph state split into three categories by retention and readership: transient, checkpoint, and persistent.
State Machine Compilation
Topology (7 nodes, 5 edges) defined in 2.4.1. Key implementation detail: checkpointer injected at graph compile time, not runtime — interrupt/resume become inherent graph properties. Routing relies on transient state fields, not node return values, because graph state must be fully serializable.
Checkpoint Persistence and Stateless Engine
Checkpoints compress runtime state to a byte sequence saved by the Go service; the engine itself stores zero session state. This makes the engine restartable and horizontally scalable — any instance can take over any suspended session. Only the latest checkpoint is kept, but pending writes must be fully exported to control blob size. Externalizing state is the prerequisite for scale and failover; any in-process persistence would break this property.
Suspend-resume follows four ordered steps: (1) engine exports checkpoint blob, (2) service persists blob, (3) service returns resume token, (4) engine imports blob and continues.
Skill Metadata Loading and Injection
At startup, Skill directory is scanned; every spec registers into the capability manifest, injected at two moments: into Planner/Reactor/ClarifyAgent for planning/judgment, and into SkillExecutor for dispatch.
Five-Layer Skill Execution Safety Pipeline
Skills are the sole gateway to real systems, so safety is paramount. The pipeline includes intent consistency verification: if a step explicitly names a Skill but the executed Skill set has empty intersection with the named set, the step fails. This catches cases where the model misinterprets ttfa_painting as coord_mesh — the latter runs without error and returns success, but the actual intent was violated. No-error and completed-expected-work are independent judgments.
Service Chain and End-to-End Presentation
Callback Validation and Three-Layer State Machines
Every engine event passes six validation checks before DB write. Validated events drive run, session, conversation state machines along terminal/non-terminal paths.
Background Compensation
When normal event flow breaks (process crash, network partition, user unresponsive), sweepsvc provides fallback compensation.
Engine Service Layer and E2E Flow
Engine exposes a stable interface: submit-and-return, async progression, full audit trail. All exceptions off the worker's main path are caught and emit error + terminal(failed) events, guaranteeing frontend receives a definitive end signal. Multi-vendor LLM switching combined with 300-second timeout prevents model single-point dependency from becoming a single point of failure.
Frontend Presentation
Frontend renders the event stream, showing step-by-step progress, tool calls, and final summary.
Design Boundaries, Trade-offs, and Evolution
Counter-Arguments
Three trade-offs stem from one principle: map data is a high-responsibility domain where errors cause user-experience accidents; controllability comes first, automation second.
Capability Placement Criteria
Decide agent vs Skill by three questions. Example: turn-restriction verification action is deterministic → Skill; whether to verify and whether to modify after verification is judgment → agent. User context (SD vs LD network) determines data level first.
Dual-Layer Evaluation and Iteration Loop
Two metric dimensions continuously feed back, forming an iteration loop that improves business fit.
Three-Phase Evolution Roadmap
Skill atomization (completed)
Chain automation (completed)
Full-chain automation (in progress)
Current system has completed the first two phases' main skeleton.
Case Studies
Case 1: External User Feedback Analysis
Old flow : five steps across disparate platforms, manual transcription between steps. New flow : single entry, three continuous steps with automatic intermediate passing, minutes end-to-end.
Case 2: National Road Database Analysis
Business needs: external support requires network mileage and risk scenario counts; internal analysis needs new element production stats and quality issue extraction. Both rely on direct national database queries. Examples: single-table statistics ("national count of toll-free toll stations") and cross-table joins ("intersections with both left-turn arrow and no-left-turn restriction") executed automatically.
Value Summary and Limitations
Quantified Benefits
Per-case handling time compressed from 20-30 minutes to 3-5 minutes → 6-10x efficiency gain .
Feedback pool issues (lane direction errors, turn-restriction redundancy, missing elements) auto-triaged; human effort shifts to conclusion review.
Current Capability Boundaries
Visual consistency between base map and real-world photos not yet covered (manual).
HITL trades efficiency ceiling for controllability.
Skill ecosystem still under construction; some domain capabilities missing.
Evaluation primarily case-level validation.
These boundaries align with architectural evolution directions.
Future Outlook
Four parallel evolution tracks: capability upgrade, harness loop, platform reuse, and new agent-driven data production mode. All share the existing orchestration layer, Skill invocation contract, and event audit trail as foundation.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Baidu Maps Tech Team
Want to see the Baidu Maps team's technical insights, learn how top engineers tackle tough problems, or join the team? Follow the Baidu Maps Tech Team to get the answers you need.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
