How China Telecom Achieved Clear WeChat Voice Calls at Just 0.266 kbps

The article analyzes China Telecom's AI Flow technology that compresses voice and video to record‑low bitrates—down to 0.266 kbps for clear calls—by replacing raw bitstreams with AI‑generated token streams, detailing the underlying theory, benchmark comparisons, and real‑world deployments across sky, air, ground, and sea scenarios.

Machine Heart
Machine Heart
Machine Heart
How China Telecom Achieved Clear WeChat Voice Calls at Just 0.266 kbps

Problem: Poor Voice Quality in Low‑Bandwidth Scenarios

In subways, underground garages, or crowded venues, weak network conditions cause WeChat voice calls to become choppy, garbled, and frustrating, leading many to assume that bad network inevitably means unintelligible audio.

WAIC Demo: 0.266 kbps Voice Call

At the 2026 WAIC China Telecom booth, two phones in a call displayed local‑network IPs and delivered crystal‑clear audio despite a displayed bitrate of 0.266 kbps —a world‑record low for voice transmission.

What Is Bitrate and Why It Matters

Bitrate (kbps) indicates how many kilobits are transmitted per second; higher rates preserve more audio detail, while lower rates force the codec to discard details, resulting in muffled or broken speech. Traditional dynamic bitrate systems raise the rate when the network is good and drop it when the network degrades.

Typical VoLTE high‑definition calls require about 15 kbps, whereas AI Flow pushes the same voice quality down to 0.266 kbps—roughly 1/500 of conventional 128 kbps codecs.

Generative Transmission: From Bits to Tokens

AI Flow’s core innovation, called Generative Transmission , follows the principle “use computation to replace bandwidth.” The sender uses a multimodal large model to “read” the speech, extracting core semantics, speaker identity, and features into a compact token stream. The receiver then uses a generative model to reconstruct the audio from these tokens, turning the transmission object from a raw bitstream into a lightweight token stream.

Underlying Theory: Signal‑Computation‑Storage (信容理论)

In 2002, Li Xuelong’s team proposed the “Signal‑Computation‑Storage” theory, showing that computation, communication, and storage can be folded into each other. The later “Positive‑Stimulus Noise” concept redefines noise as potentially useful when correlated with the task, allowing AI to retain useful cues even at ultra‑low bitrates.

Extending to Video: Token‑Based 1080p at <100 kbps

For video, AI Flow extracts semantic and dynamic features from the 1080p 24 FPS stream, encodes them into tokens, and transmits them. The receiver’s generative model reconstructs the video, achieving stable transmission at under 100 kbps—about 3 % of the ~3.6 Mbps required by traditional codecs.

World Model for 3‑D Reconstruction

By coupling token streams with a world model, AI Flow can transmit an entire soccer match’s spatiotemporal relationships, allowing the receiver to reconstruct a navigable 3‑D scene where viewers can switch perspectives, including first‑person goalkeeper view.

Four Deployment Scenarios: Sky, Air, Ground, Sea

Sky : Deploy generative models on satellites, turning the satellite into a compute node that decodes tokens, while ground terminals only send lightweight tokens.

Air : Small UAVs equipped with AI Flow boxes and miniature satellite links can stream 720p video at ~60 kbps, far below the ~2 Mbps needed by H.264.

Ground : Robots achieve end‑to‑end latency of 20‑50 ms (down to 20 ms) over satellite links, enabling nationwide remote operation and shared semantic perception across heterogeneous robots.

Sea : Using high‑frequency acoustic links, AI Flow transmits real‑time 720p video over >10 m underwater at low bandwidth, supporting inspection, ecological monitoring, and rescue.

Real‑World Applications

Emergency rescue: An app‑based AI Flow system sends a single image at 9.6 kbps and 720p video at ~110 kbps over narrow satellite links, already deployed in 31 Chinese provinces with over a thousand units.

Mass video storage: Token‑based encoding reduces storage to ~1/40 of traditional methods, extending a week‑long storage capacity to nearly a year; tens of thousands of cameras have been upgraded on the Tianyi Vision network.

Privacy‑preserving monitoring: Semantic token filtering removes sensitive visual data while preserving behavior alerts, enabling privacy‑first deployments in elder‑care and campus safety for ~400 million endpoints.

Impact and Outlook

AI Flow demonstrates that communication is no longer merely a transport tool but a condition for intelligent emergence. By converting bits into semantic tokens, it breaks bandwidth ceilings and allows AI, agents, and robots to operate wherever connectivity was previously impossible.

As Li Xuelong remarked, “Intelligent emergence stems from connection and interaction.” The technology embodies this by turning the act of connecting devices into the very substrate of future AI growth.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

multimodal AIToken CompressionChina TelecomAI FlowLow BandwidthGenerative Transmission
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.