How China Telecom Achieved Clear WeChat Voice Calls at Just 0.266 kbps
The article analyzes China Telecom's AI Flow technology that compresses voice and video to record‑low bitrates—down to 0.266 kbps for clear calls—by replacing raw bitstreams with AI‑generated token streams, detailing the underlying theory, benchmark comparisons, and real‑world deployments across sky, air, ground, and sea scenarios.
Problem: Poor Voice Quality in Low‑Bandwidth Scenarios
In subways, underground garages, or crowded venues, weak network conditions cause WeChat voice calls to become choppy, garbled, and frustrating, leading many to assume that bad network inevitably means unintelligible audio.
WAIC Demo: 0.266 kbps Voice Call
At the 2026 WAIC China Telecom booth, two phones in a call displayed local‑network IPs and delivered crystal‑clear audio despite a displayed bitrate of 0.266 kbps —a world‑record low for voice transmission.
What Is Bitrate and Why It Matters
Bitrate (kbps) indicates how many kilobits are transmitted per second; higher rates preserve more audio detail, while lower rates force the codec to discard details, resulting in muffled or broken speech. Traditional dynamic bitrate systems raise the rate when the network is good and drop it when the network degrades.
Typical VoLTE high‑definition calls require about 15 kbps, whereas AI Flow pushes the same voice quality down to 0.266 kbps—roughly 1/500 of conventional 128 kbps codecs.
Generative Transmission: From Bits to Tokens
AI Flow’s core innovation, called Generative Transmission , follows the principle “use computation to replace bandwidth.” The sender uses a multimodal large model to “read” the speech, extracting core semantics, speaker identity, and features into a compact token stream. The receiver then uses a generative model to reconstruct the audio from these tokens, turning the transmission object from a raw bitstream into a lightweight token stream.
Underlying Theory: Signal‑Computation‑Storage (信容理论)
In 2002, Li Xuelong’s team proposed the “Signal‑Computation‑Storage” theory, showing that computation, communication, and storage can be folded into each other. The later “Positive‑Stimulus Noise” concept redefines noise as potentially useful when correlated with the task, allowing AI to retain useful cues even at ultra‑low bitrates.
Extending to Video: Token‑Based 1080p at <100 kbps
For video, AI Flow extracts semantic and dynamic features from the 1080p 24 FPS stream, encodes them into tokens, and transmits them. The receiver’s generative model reconstructs the video, achieving stable transmission at under 100 kbps—about 3 % of the ~3.6 Mbps required by traditional codecs.
World Model for 3‑D Reconstruction
By coupling token streams with a world model, AI Flow can transmit an entire soccer match’s spatiotemporal relationships, allowing the receiver to reconstruct a navigable 3‑D scene where viewers can switch perspectives, including first‑person goalkeeper view.
Four Deployment Scenarios: Sky, Air, Ground, Sea
Sky : Deploy generative models on satellites, turning the satellite into a compute node that decodes tokens, while ground terminals only send lightweight tokens.
Air : Small UAVs equipped with AI Flow boxes and miniature satellite links can stream 720p video at ~60 kbps, far below the ~2 Mbps needed by H.264.
Ground : Robots achieve end‑to‑end latency of 20‑50 ms (down to 20 ms) over satellite links, enabling nationwide remote operation and shared semantic perception across heterogeneous robots.
Sea : Using high‑frequency acoustic links, AI Flow transmits real‑time 720p video over >10 m underwater at low bandwidth, supporting inspection, ecological monitoring, and rescue.
Real‑World Applications
Emergency rescue: An app‑based AI Flow system sends a single image at 9.6 kbps and 720p video at ~110 kbps over narrow satellite links, already deployed in 31 Chinese provinces with over a thousand units.
Mass video storage: Token‑based encoding reduces storage to ~1/40 of traditional methods, extending a week‑long storage capacity to nearly a year; tens of thousands of cameras have been upgraded on the Tianyi Vision network.
Privacy‑preserving monitoring: Semantic token filtering removes sensitive visual data while preserving behavior alerts, enabling privacy‑first deployments in elder‑care and campus safety for ~400 million endpoints.
Impact and Outlook
AI Flow demonstrates that communication is no longer merely a transport tool but a condition for intelligent emergence. By converting bits into semantic tokens, it breaks bandwidth ceilings and allows AI, agents, and robots to operate wherever connectivity was previously impossible.
As Li Xuelong remarked, “Intelligent emergence stems from connection and interaction.” The technology embodies this by turning the act of connecting devices into the very substrate of future AI growth.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
