openJiuwen WorkSwarm: Full-Duplex Multimodal Agent for Real-Time Chat & Background Tasks

openJiuwen WorkSwarm introduces a full-duplex multimodal AI agent platform where a real-time model handles live audio/video interaction while a Core Agent executes complex tasks in the background, allowing users to interrupt, reprioritize, and continue chatting without waiting for tool calls to finish.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
openJiuwen WorkSwarm: Full-Duplex Multimodal Agent for Real-Time Chat & Background Tasks

AI assistants are entering a new working rhythm: conversing while invoking tools to get things done. Recent coverage of Gemini Live updates highlighted background tool calling during real-time interaction — the assistant can acknowledge a request, then continue processing time‑consuming work. This raises a practical question: when a tool task is running, can the AI still respond to new user input?

Understanding Full‑Duplex Interaction

Traditional voice assistants operate like walkie‑talkies: the user finishes, then the AI replies; any interruption is ignored or breaks the flow. Full‑duplex is more like a phone call — both sides can speak and listen simultaneously. For the user, the most direct change is the elimination of waiting: no need to wait for the AI to finish speaking, and no need to adapt to its turn‑taking pace.

WorkSwarm’s Dual‑Path Architecture

openJiuwen is an open‑source AI agent platform jointly built by Huawei’s 2012 Lab, Huawei Cloud, Terminal, Computing, and Compute Pioneer teams together with universities and enterprise developers. WorkSwarm is its open‑source unified workspace for office and coding. It merges real‑time audio/video interaction with Core Agent execution into a single task: users can show the live camera view, interrupt at any moment, and hand off further processing to the background Core Agent.

Interruption Handling

Human conversation rarely follows strict turn‑taking. Users may interject with “Skip the background, just give the conclusion.” WorkSwarm detects new valid speech via voice activity detection and a noise gate to avoid false triggers, stops the current playback, and processes the new instruction while preserving already generated text in the task record. A demo shows the agent reading the Yueyang Tower Preface ; the user interrupts to request reading the Lantingji Xu instead, and the switch occurs naturally and quickly.

Interruption affects only the current dialogue layer. Tasks already delegated to the Core Agent run independently; cancelling them requires a manual stop in the task control panel. The system simultaneously manages two layers: the conversation layer accepts new requests, while the task layer continues executing ongoing work.

Background Task Execution with Live Status

A simple visual question (“What brand is this?”) can evolve into a multi‑step task (“Search its info and compile a document”). WorkSwarm routes this through two paths: the real‑time model handles perception and immediate response; the Core Agent decomposes the task, calls tools (search, analysis, file generation), and pushes progress forward. While search runs in the background, the user can still ask about the live view or interrupt the spoken summary.

The UI continuously displays task status, tool invocation logs, execution results, and generated files — not just a spinning loader. In a recorded demo, a task to read an algorithm Word document runs visibly in the top‑right task queue while the user continues chatting. Upon completion, the system first writes a textual summary into the chat box, then waits for the current dialogue turn to end before verbally reporting the result, extracting only key points rather than reading the entire document.

Task Queue Management and Decoupled Connections

If a second request arrives before the first finishes, WorkSwarm’s full‑duplex task queue provides waiting slots. Users can reorder tasks, promote one to run next, or stop waiting/running tasks. Dialogue continues while task states remain clearly visible.

Audio/video connections and tool tasks are managed separately. Closing the full‑duplex session does not terminate Core Agent jobs; results, tool outputs, and files return to the original conversation when the user reconnects. For example, after launching a brand query via camera, the user can end the call and later return to view the gathered materials. This gives on‑the‑spot instructions a traceable execution trail.

Deployment and Model Configuration

WorkSwarm offers one‑click installers for Windows, macOS, and HarmonyOS. Model and voice channels are configurable in settings. Two protocols are supported:

Qwen Omni realtime websocket — accepts the official base URL or a locally deployed endpoint.

JoyAI — a three‑model pipeline (ASR, LLM, TTS) that allows using official APIs or third‑party ASR/TTS models.

Additionally, the platform supports deploying full‑duplex multimodal models on Huawei Ascend NPUs via the vLLM‑Omni inference framework on Huawei Cloud ModelArts, then connecting them to WorkSwarm’s multimodal full‑duplex capability.

Key Takeaways

WorkSwarm’s full‑duplex multimodal capability is more than “listening while speaking.” It enables true parallelism between real‑time conversation and Core Agent background execution: users interact naturally in the foreground — interrupting, follow‑up questioning, adding requirements — while complex tasks proceed at their own pace in the background, fetching data, invoking tools, and performing analysis. This “conversation never pauses, tasks never stop” experience has moved beyond demos into a usable collaboration mode. When AI no longer forces humans into rigid turns but collaborates like a colleague who chats while working, the next rhythm of human‑machine teamwork begins.

Technical references:

GitHub: https://github.com/openJiuwen-ai AtomGit: https://atomgit.com/openJiuwen Agentic Hub: https://agentichub.openjiuwen.com vLLM‑Omni Ascend deployment example:

https://github.com/vllm-project/vllm-omni/tree/main/examples/online_serving/minicpmo

Huawei Cloud ModelArts docs:

https://support.huaweicloud.com/modelarts/index.html
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

real-time interactionfull-duplexAscend NPUopenJiuwenvLLM-Omnibackground task executionmultimodal AI agenttask queue management
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.