Why Multimodal AI Is the Next Battlefield After Coding – Insights from SenseTime’s Lin Dahua

In a WAIC interview, SenseTime’s chief scientist Lin Dahua explains why multimodal AI, embodied in the native‑unified NEO‑unify architecture and the commercial‑grade SenseNova U1 Pro, is poised to surpass coding as the next competitive frontier, highlighting technical challenges, data efficiency, design‑focused benchmarks, and a 70% delivery‑rate claim.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Why Multimodal AI Is the Next Battlefield After Coding – Insights from SenseTime’s Lin Dahua

During the WAIC 2026 conference, SenseTime’s chief scientist Lin Dahua was asked where the next decisive battleground would be after large language models. He illustrated the limitation of pure text by pointing to a potted plant and asking the interviewee to describe every leaf’s exact position, emphasizing that visual perception cannot be fully captured by language alone.

Lin argues that any scenario requiring understanding of physical space—high‑level aesthetics, urban layout, autonomous driving, embodied intelligence—demands true visual capability. From this first‑principle, SenseTime bets on multimodal AI as the next frontier.

Over the past year, major players such as OpenAI, Kimi, DeepSeek, ByteDance, and Alibaba have heavily invested in multimodal research. Unlike many, SenseTime pursues a “hard‑core” path: a bottom‑up fusion of language and vision through a native unified architecture called NEO‑unify . In March 2024 the company released the open‑source multimodal model SenseNova U1 , and at the WAIC forum it unveiled the commercial‑grade flagship SenseNova U1 Pro , which claims native 8K image quality and positions itself against top models like GPT‑Image‑2.

The NEO‑unify design removes the conventional visual encoder and variational auto‑encoder, allowing the model to learn directly from pixels and text. Its backbone is a self‑developed Mixture‑of‑Transformers (MoT) that interleaves language and vision at every layer, a reconstruction comparable to the impact of the original Transformer on LLMs. Remarkably, the architecture achieves top‑tier visual perception and reasoning with only 390 million image‑text pairs—about one‑tenth of the data required by comparable industry models.

U1 Pro’s performance is highlighted by a reported 70% commercial delivery rate, meaning generated designs need little to no post‑editing to meet professional standards. This metric surpasses the 60% baseline for commercial viability and matches only two multimodal models globally (the other being OpenAI’s GPT‑Image‑2). The model achieves thousand‑level error control for text, layout, color, and element placement.

To reach this level, SenseTime built a massive multimodal dataset exceeding a billion design‑level examples, many of which contain thousands of words describing a single graphic. Training proceeds in three stages: massive ingestion of raw image‑text pairs, heavy reliance on synthetic data generated by an internal agent pipeline, and a final fine‑tuning phase on a curated “golden” subset of tens of millions of examples.

U1 Pro also incorporates a “visual decomposition thinking chain” that first breaks down a high‑quality poster into visual components, infers the designer’s layout logic and color scheme, and then learns this reasoning process. Generation follows a multi‑round intelligent agent loop: (1) textual requirement is decomposed into a layout plan, (2) the model produces the full image, (3) a global self‑check identifies textual errors, layout flaws, or aesthetic issues and iteratively refines the output until it meets the delivery threshold.

Quality assurance is reinforced by a 200‑person “human aesthetic team” of art‑school students and professional designers who continuously critique and score U1 Pro’s outputs, akin to a rigorous apprenticeship that teaches the model what true high‑level design feels like.

Beyond static design, SenseTime is extending the multimodal foundation toward three‑dimensional understanding. The upcoming U2 model aims to integrate video perception and generation, enabling tasks such as turning a sketch into a full 3‑D building. The company also released a vision‑only model, SenseNova‑Visio n , which demonstrates native visual capabilities that surpass specialized vision models on tasks like detection, segmentation, depth prediction, and 3D reconstruction.

Lin envisions the ultimate “world model” as a unified system where vision, language, and dozens of sensor modalities share a common core representation, allowing seamless reasoning across physical and social dimensions. He acknowledges the exponential growth in data and compute required for such a model and proposes a modular approach: decouple physical attributes into specialized sub‑models that can be composed for complex inference.

In the interview’s closing, Lin declares the next technical summit his team will tackle: building AGI that truly understands and acts in the physical world.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AIlarge language modelsmultimodaldesign AISenseNovaNEO-unify
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.