Voice Agent Architecture: Full-Duplex Frontend + Async Backend Delegation
This article explains a voice agent architecture that separates real-time full-duplex frontend interaction from asynchronous backend computation using delegation IDs, detailing the two-timeline mechanism, three-layer responsibility split, task state machine, and evaluation strategies to avoid stale responses and interruption mishandling.
01 | Why a Single Pipeline Blocks
The classic cascade ASR → LLM → TTS assumes one speaker finishes before the system responds. Real conversations involve interruptions, corrections, and mid‑sentence requests (e.g., “wait, check that”). Any design that waits for a full turn exposes latency gaps.
Tool calls often take seconds to tens of seconds, while spoken feedback needs sub‑100 ms response. Binding them together forces either fake acknowledgments (“okay, hold on”) with no actual work, or completed work that arrives after the user has changed topic.
Core principle: interaction latency budget and task latency budget are different orders of magnitude, so they cannot share a single blocking path.
02 | Voice Frontend: What Full‑Duplex Actually Handles
Full‑duplex means the system continues receiving user audio while playing its own speech, and can classify incoming audio as back‑channel or true interruption.
The frontend’s concrete responsibilities:
Endpointing & turn‑taking: has the user finished or just paused?
Barge‑in handling: how to instantly converge output audio when the user cuts in.
Prosody & rhythm: when to pause, back‑channel, or continue speaking.
Lightweight understanding: does this utterance need delegation? Keyword/intent hit?
Boundary: the frontend may produce transcripts and spoken replies, but must not own long‑running reasoning, complex tool orchestration, or permission approvals — those would break real‑time guarantees.
03 | Core Mechanism: Two Timelines
After splitting, two clocks run simultaneously:
Voice timeline: listen, speak, interrupt, re‑inject spoken results.
Task timeline: queue, execute, succeed/fail, expire or cancel.
Delegation is not “handing the microphone to the backend.” The frontend emits a task with a delegation ID, keeps the session alive, and later receives the conclusion under the same ID to decide how to verbalize it.
Key invariant: an interruption changes the voice timeline but does not automatically cancel the task timeline. Cancellation must be explicitly modeled.
04 | Three‑Layer Responsibility Split
To avoid a “do‑everything voice model” swallowing business logic, split into:
Voice Frontend Layer: real‑time transcription, interruption, spoken generation.
Delegation Bus: session context slicing, task creation, result re‑injection, timeout policies.
Backend Layer: stronger reasoning models, tools, CRM, knowledge bases, approval flows.
Two common delegation shapes: managed backend (session directly specifies a reasoning model) and client delegation (your own Agent/service picks up the task). Mechanism is identical; difference lies in who executes and who controls permissions.
05 | Task State Machine: Handling Late Results
Without a state machine, dual timelines inevitably race. Minimum states: Idle → In‑Session → Delegated → Waiting‑Result → Returned; plus side branches: Expired / Cancelled.
Three frequent pitfalls:
User has switched topic, but the old query result is still spoken aloud.
Interruption stops TTS, yet background writes continue (“speech stopped, work didn’t”).
Multiple concurrent delegations return out of order, causing contradictory spoken output.
Practical strategies: before re‑injection, verify “does this delegation still match the current topic?”; expired results go only to logs or a light toast; tools with side effects default to confirmation or rollback capability.
06 | What This Means for R&D
This is not a TTS upgrade — it’s a concurrency model shift. You must observe simultaneously: session events, delegation events, tool durations, interruption timestamps. Correlate everything with session ID + delegation ID.
Evaluation must also split: frontend metrics — first‑packet latency, interruption recovery, turn naturalness; backend metrics — tool success rate, correctness, tail latency. Merging into a single “sounds good” score hides failures.
Selection reminder: cascade architectures aren’t obsolete — logs are clear, compliance is easier. The new paradigm fits high‑interaction, high‑tool scenarios; both can coexist, chosen by scenario, not belief.
Mental Model Summary
Voice frontend guarantees “presence”; backend delegation guarantees “execution”; delegation ID guarantees both sides stay aligned on the conversation. Master these three and you grasp the mechanism’s backbone — everything else is protocol detail and product strategy.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
