WorkSwarm: Full-Duplex Multimodal AI Agent Enables Chat-Like Interaction With Background Task Execution

openJiuwen WorkSwarm introduces full-duplex multimodal interaction where a real-time model handles audio/video while a Core Agent runs background tasks, allowing users to interrupt, delegate work, and manage a persistent task queue across Windows, Mac, and HarmonyOS with support for JoyAI, Qwen Omni, and vLLM-Omni on Ascend NPU.

Machine Heart
Machine Heart
Machine Heart
WorkSwarm: Full-Duplex Multimodal AI Agent Enables Chat-Like Interaction With Background Task Execution

Overview

openJiuwen is an open-source AI Agent platform jointly built by Huawei's 2012 Lab, Huawei Cloud, terminal, computing, and compute-power teams together with universities and developers. WorkSwarm is its open-source unified workspace for office and coding that combines real-time audio/video interaction with Core Agent execution in a single task.

Full-Duplex vs. Half-Duplex Interaction

Traditional voice assistants operate like walkie-talkies (half-duplex): the user finishes speaking, then the AI responds; interruptions are ignored or break the flow. Full-duplex works like a phone call: the AI can listen and speak simultaneously, and the user can interrupt or add input at any moment. This compresses the waiting experience — users no longer need to wait for the AI to finish or adapt to its rhythm.

Interrupt Handling

When the system detects new valid user speech via voice activity detection (VAD) and a noise gate, it stops the current playback and begins processing the new instruction. Previously generated text remains in the task record. For example, during a reading of Yueyang Tower Record , a user can say "read Lantingji Xu " and the agent switches naturally.

Interruptions affect only the current dialogue layer. Tasks already delegated to the Core Agent run independently; cancelling them requires manual action in the task control panel.

Dual-Path Architecture: Real-Time Model + Core Agent

WorkSwarm separates concerns into two parallel paths:

Real-time model : handles screen viewing, listening, and immediate responses.

Core Agent : decomposes tasks, calls tools (search, analysis, file generation), and drives long-running work.

While the Core Agent searches in the background, the conversation continues in the foreground. Users can still ask about the screen, interrupt playback, or add requirements.

Task Queue Management

The full-duplex task queue gives each subsequent request a waiting slot. Users can reorder tasks, promote a task to run next, or stop waiting/running tasks. The UI shows task status, tool calls, execution results, and file artifacts — not just a spinning loader.

Audio/video connections and tool tasks are managed independently. Closing the real-time session does not terminate Core Agent tasks; results (text, tool outputs, files) return to the original conversation when the user reconnects. For instance, after launching a brand query via camera, the user can end the video session and later retrieve the compiled research.

Model Protocol Support & Deployment

WorkSwarm supports two protocol stacks:

Qwen Omni realtime websocket : accepts official base_url or a local deployment address.

JoyAI : a three-model pipeline (ASR, LLM, TTS) that can use official APIs or third-party ASR/TTS models.

Additionally, full-duplex multimodal models can be deployed on Ascend NPUs via Huawei Cloud ModelArts using the vLLM-Omni inference framework (

https://github.com/vllm-project/vllm-omni/tree/main/examples/online_serving/minicpmo

).

Installation & References

One-click installers are provided for Windows, Mac, and HarmonyOS. Key resources:

Official site: https://openjiuwen.com Download: https://www.openjiuwen.com/download GitHub: https://github.com/openJiuwen-ai AtomGit: https://atomgit.com/openJiuwen Agentic Hub: https://agentichub.openjiuwen.com ModelArts deployment guide:

https://support.huaweicloud.com/modelarts/index.html
Complex agent tasks run in the background while dialogue continues in the foreground. Real-time models and Core Agents now operate at two different rhythms, delivering a new collaborative experience where conversation never pauses and tasks never stop.
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Open sourceAI AgentMultimodalReal-Time InteractionTask QueueHuaweiFull-DuplexAscend NPU
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.