Edge AI Token Factories: Banma Smart's AutoOmni 2.0 Enables Deep AI Usage in Cars

At the 2026 Yunqi Conference, Banma Smart unveiled AutoOmni 2.0-23B-A3B, a sparse MoE edge model that achieves near-cloud performance on vehicle-grade chips, addressing memory, latency, and concurrency constraints while enabling privacy-preserving, low-cost AI with certified safety frameworks for mass production.

Machine Heart
Machine Heart
Machine Heart
Edge AI Token Factories: Banma Smart's AutoOmni 2.0 Enables Deep AI Usage in Cars

Shift from Cloud to Edge: The New AI Battleground

The 2026 Yunqi Conference highlighted a pivotal industry shift: AI competition is moving from cloud model capability to edge deployment — whether models can run stably and securely under limited memory, compute, and network conditions. Banma Smart (斑马智能), an Alibaba Group subsidiary, showcased its AutoOmni series, a full-modal on-device large model family for intelligent cockpits. The new flagship AutoOmni 2.0-23B-A3B is a sparse Mixture-of-Experts (MoE) model with 23B total parameters and only 3B activated per inference, designed to run on automotive-grade SoCs like the Snapdragon 8397.

Three Walls and Two Industry Constraints

Banma CTO Si Luo (司罗) identifies three hard walls for edge AI:

Memory Wall : Vehicle SoCs have far less RAM than servers, shared with other cockpit processes. Model viability depends on whether weights can reside in memory, not peak compute.

Latency Wall : Cockpit interactions tolerate ~100ms; active safety alerts need even less. Cloud round-trips and network jitter can consume the entire budget.

Concurrency Wall : Real cockpits run voice, vision, vehicle control, navigation, and music simultaneously. Models must stay stable under multi-task loads, not just single-task benchmarks.

Beyond these, two industry constraints dominate:

Privacy : In-cabin voice, video, and location are high-sensitivity data; full cloud upload is increasingly untenable. Si Luo calls the car a "very important semi-closed scene requiring privacy protection."

Cost : Agent-era token consumption is a recurring expense. Offloading everything to the cloud forces OEMs to pay long-term bills for every vehicle in use.

Si Luo argues edge models unlock the value of already-paid hardware: "Turn on-device capability into a true on-device token factory," drastically cutting uplink traffic and cloud token costs.

Sparse Architecture: The Edge Solution

MoE matters more on edge than cloud. In the cloud, MoE improves training/inference economics; on edge, it decouples total parameters (capability ceiling) from activated parameters (per-inference memory/compute), enabling near mid/large-model capability within tight physical resources. Banma's engineering targets are summarized as 多、快、好、省 (multi, fast, good, thrifty):

Multi : 8+ concurrent tasks stable.

Fast : 5–6× inference speedup.

Good : >99% quantization fidelity.

Thrifty : 50% runtime memory reduction.

The last two are decisive: aggressive quantization typically sacrifices accuracy; fidelity measures what remains. Halving memory directly determines co-existence with other cockpit processes.

From 4B Dense to 23B-A3B MoE: Raising the Ceiling

Banma's previous edge mainstay was a 4B dense model. AutoOmni-4B delivered 20% average cockpit-scenario improvement over general LLMs: +22% text understanding (vehicle control, navigation, entertainment), +17% visual scene understanding, +122% audio/emotion understanding, while matching top general LLMs on broad tasks.

AutoOmni 2.0-23B-A3B pushes further. Si Luo states: "This 3B-activated MoE model matches 10× parameter cloud models on ordinary tasks, and reaches 80–90% of 10× cloud models on complex tasks." Evaluation covers four complex task categories; edge performance approaches cloud across the board. A live demo showed continuous multi-intent commands: adjust seat, start front massage, enable full-car ventilation, then multi-turn navigation with dynamic waypoint insertion and re-routing.

Notably, Si Luo admitted that three weeks prior he deemed a demo impossible — edge model progress is outpacing even practitioners' expectations.

Edge-Cloud Collaboration: Big Brain / Small Brain Division

Edge strength doesn't retire the cloud; it demands clear division. Banma splits understanding & planning (big brain) to AutoOmni on device, and intent decomposition, task routing, controlled execution (small brain) to AutoClaw, System Agent, Harness .

The edge-cloud framework decides: which tasks run cloud vs. edge, and how data is partitioned (e.g., raw privacy data stays local; summarized memories may sync to cloud). Under this, on-device Agent Groups (navigation, media, proactive vision, GUI) are orchestrated by SystemAgent for domain routing, rejection (拒识), proactive engagement, memory signal extraction, and edge-cloud arbitration (both sides produce results; system adjudicates). Memory forms a dedicated layer with real-time extraction, personal info registration, and RAG-backed storage via SQLite.

Rejection (拒识) is critical: a demo showed the system listening to passenger chatter but responding only when addressed ("close the window"). Si Luo: "The hard part of full-time wake-free isn't hearing — it's knowing when not to answer."

Cloud handles knowledge and complex reasoning. Cai Ming (蔡明), Chief Product Officer, summarizes three connectivities: (1) edge-cloud compute/service connectivity (YuanShen AI core), (2) cross-device connectivity around the user (full understanding of "my world"), (3) user-service connectivity (cross-scenario, cross-device, cross-app integration).

Extending to Intelligent Driving

Banma is applying the same framework to ADAS/AD. Phase 1: cockpit NLU/planning calls driving-system APIs (e.g., activate_noa(destination), lane_change(direction)) via a safety boundary built from rules + prompts. Phase 2: deeper model-level fusion where cockpit VLM/VLA outputs directly enhance driving perception, planning, and long-tail handling. The safety boundary remains a distinct layer: model suggests "keep distance from truck"; system ensures that understanding never becomes a dangerous lane change.

Safety & Compliance: Three-Layer Defense + Cloud Ops

Edge deployment raises two safety questions: internal control (can the system contain the model?) and external certification (can it enter regulated markets?). Si Luo: "With edge capability, edge safety becomes a newly critical function."

Banma's vehicle-grade edge AI safety system has three defense-in-depth layers:

Environment Layer : trusted model loading, inference sandbox, resource isolation, mandatory access control, runtime intrusion detection — ensuring an isolated, auditable runtime.

Framework Layer : identity propagation & permission persistence, MCP security protocol, context/memory isolation, risk monitoring & full-lifecycle audit — constraining every Agent tool/service call.

Content Layer : on-device lightweight safety model for generation-time compliance checks.

Above these, Cloud Security Operations trains/updates safety models, operates content policies, and provides threat perception via a trusted edge-cloud channel. Policies and models iterate in cloud, deploy via trusted channel, but every verification still runs locally — so safety never depends on network availability.

Results: >98% jailbreak attack recall. Edge safety model matches industry Guard models in accuracy, with 4× throughput and <3% system performance impact. Local verification means sensitive data never leaves the vehicle and safety doesn't degrade with network conditions.

Certification as Entry Ticket for Global Markets

Chinese cockpit UX leads globally (1–2 generations ahead per Si Luo), with BMW iX3, VW, Porsche adopting Chinese tech stacks. But experience ≠ market access. Europe mandates GDPR, UN R155 (cybersecurity), UN R156 (software updates), UN R171 (driver control assistance), plus upcoming Chinese national standards in 2027.

Banma developed Banma Safety Linux (ASIL-B functional safety certified) and achieved ASPICE 4.0 CL2 (German auto industry standard, 20+ OEMs, significantly strengthened for AI/ML) in January 2026. In automotive, certification is a prerequisite: a model can ship only if its software engineering process is third-party proven controllable, irrespective of leaderboard rank. Owning OS, edge model, execution framework, chip inference, and cloud ecosystem — all certified and in mass production — is Banma's moat.

This constitutes a concrete sample of edge technology sovereignty : the full trust chain from kernel to application is self-controlled and externally verified.

Car as the First Proving Ground for Embodied AI

The eight vehicles outside the venue (from Denza, Xingji, Hongqi, IM, MG, etc.) embody one judgment: "Cloud decides how high China's AI can reach; edge decides whether AI is allowed to act in the real world." Cai Ming frames token as the fuel of the intent economy — the new currency for fulfilling user intent. For that currency to circulate in cars, edge must close the loop on understanding, execution, and safety locally.

Edge is increasingly treated as infrastructure. If cloud competition sets China's AI height, edge sets its depth and safety boundary, and technology sovereignty ultimately lands in mass-producible terminals. Si Luo: "Intelligent vehicles are the largest proving ground and scale stress-test platform for embodied AI." They simultaneously impose memory, latency, concurrency, privacy, and cost constraints — which will reappear with different weights on phones, glasses, wearables, and robots. Banma calls this extrapolation from AutoOmni to EdgeOmni : from a single car's edge brain to multi-terminal intelligence around one user, from solitary to collective intelligence. Whoever solves the system engineering puzzle on wheels first holds the migration template for all other endpoints.

The forum delivered a phased answer to that puzzle.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Edge AIMixture of ExpertsEmbodied AIAI safetyon-device AICloud-Edge CollaborationModel QuantizationAutomotive AIASPICE CertificationAutoOmni
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.