Why Multimodal AI Is the Next Battlefield After Coding, According to SenseTime’s Lin Dahua
In an interview at the 2026 WAIC conference, SenseTime chief scientist Lin Dahua explains why multimodal AI—driven by the native unified NEO‑unify architecture and embodied in the commercial‑grade SenseNova U1 Pro with a 70% delivery rate—represents the next decisive frontier beyond AI coding, highlighting technical challenges, market trends, and future research directions.
On July 18, during the WAIC 2026 conference, AI Tech Review interviewed SenseTime’s chief scientist Lin Dahua about the next decisive battlefield after large‑language‑model (LLM) coding.
When asked where the next competitive edge lies, Lin illustrated the limitation of pure language by pointing to a potted plant and asking for a precise description of each leaf’s position, emphasizing that visual perception cannot be fully captured by text.
He argued that any scenario requiring understanding of physical space—high‑level aesthetics, 3D city layouts, autonomous driving, embodied intelligence—demands genuine visual capability. From this first‑principle, Lin stated that multimodal AI is the next battlefield for SenseTime.
Why Multimodal Over Pure Coding?
Lin acknowledged that AI coding is currently the most mature commercial lane, but the ceiling of pure‑text generation is already evident. Code‑generated front‑end pages lack the “Apple‑grade” visual quality of professional designs, revealing a blind spot for text‑only models.
Industry leaders such as Anthropic are already strengthening multimodal elements (e.g., Opus 4.8), indicating that as backend code capabilities converge and price wars intensify, front‑end experience and visual interaction become the next competitive high ground. SenseTime therefore prefers to leverage its long‑standing visual expertise rather than chase short‑term coding gains.
Multimodal Architecture: From “Splicing” to Native Unification
The prevailing multimodal route has been the “splicing architecture,” which attaches an independent visual encoder (VE) to a language model, translating images into text before processing. While this approach lowers engineering barriers and speeds training, it suffers from low fusion efficiency and loss of fine‑grained visual details.
SenseTime identified these limits early and pursued a native unified multimodal architecture called NEO‑unify . This design removes the visual encoder and variational auto‑encoder, allowing the model to learn directly from pixels and text without a translation layer.
The core of NEO‑unify is a self‑developed Mixture‑of‑Transformers (MoT) framework that interleaves language and vision information at every layer, achieving a reconstruction of multimodal fusion comparable to the impact of the Transformer on LLMs.
Empirically, NEO‑unify reaches top‑tier visual perception and reasoning using only 3.9 × 10⁸ image‑text pairs—about one‑tenth of the data required by comparable industry models.
SenseNova U1 and U1 Pro: From Research to Commercial Delivery
In April, SenseTime open‑sourced the multimodal model SenseNova U1 , built on the native unified architecture, featuring an “interleaved image‑text thinking chain” that can write text and insert images within the same context, delivering native 8K‑quality outputs that rival leading models such as GPT‑Image‑2.
At the WAIC forum, SenseTime unveiled the commercial‑grade flagship SenseNova U1 Pro . Marketed as the world’s first multimodal tool with “understand‑generate‑act” as its core, U1 Pro achieves a 70% delivery rate —meaning generated designs require minimal or no post‑editing and meet professional designer standards.
Only two multimodal models surpass this benchmark: SenseNova U1 Pro and OpenAI’s GPT‑Image‑2, which is considered the industry’s de‑facto standard.
U1 Pro’s high delivery rate stems from rigorous full‑dimensional quality control: text error rates are kept at the per‑thousand level, and visual element, layout, and color errors are similarly constrained.
Visual Decomposition Thinking Chain
SenseTime introduced a “visual decomposition thinking chain” training mode that transforms the model into a “thinking designer.” During training, the model first disassembles a high‑quality poster into visual components, then reverse‑infers the designer’s layout logic and color palette, learning the underlying design reasoning rather than mere pixel patterns.
Generation follows a multi‑round intelligent agent loop: (1) decompose textual requirements into a complete design plan; (2) generate the full image using the learned visual‑language fusion; (3) perform global self‑inspection to detect textual errors, layout flaws, and aesthetic issues, iteratively refining until the output meets the delivery standard.
To polish the final aesthetic, SenseTime assembled a “human aesthetic team” of 200 art‑school students and professional designers who rigorously evaluate and score U1 Pro outputs, akin to a design studio’s iterative critique process.
Beyond 2D Design: World Models and Future Directions
Lin highlighted that U1 Pro already exhibits nascent spatial reasoning abilities, such as solving a Rubik’s cube, hinting at emerging 3D capabilities. The broader multimodal foundation also supports “embodied intelligence” and world‑model research.
SenseTime’s “1+X” strategy has spawned ecosystem ventures like the “Kaiwu” world model, which extends the visual‑language base with higher proportions of embodied data.
In July, SenseTime released another unified visual model, SenseNova‑Visio n , which positions native vision as a universal foundation model, surpassing specialized visual models across tasks like detection, segmentation, depth prediction, and 3D reconstruction.
Looking ahead, the next generation U2 aims to incorporate video understanding and generation, moving from 2D design to full 3‑D creation—potentially enabling a model to generate an entire building from a sketch.
Ultimately, Lin envisions a “world model” that integrates dozens of sensor modalities beyond vision and language, tackling physical‑world AGI challenges. He anticipates that scaling such models will demand novel strategies, such as decoupling world attributes into specialized sub‑models for collaborative reasoning.
Key Publications
Lin Dahua et al., “Towards Native Vision‑Language Primitives at Scale,” 2025.
Lin Dahua et al., “NEO‑unify: Building Native Multimodal Unified Models End to End,” 2026.
Overall, the interview underscores that delivering commercially viable multimodal AI—especially in design‑heavy scenarios like infographics, advertising, and B2B visual content—requires deep visual‑language integration, massive high‑quality multimodal data, and rigorous quality‑control pipelines, positioning multimodal AI as the next strategic frontier after coding.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
