Breaking Modality Barriers with an 11B Multimodal Scientific Model “ShenZhen”
The 11‑billion‑parameter multimodal scientific foundation model “ShenZhen” unifies DNA, RNA, protein, small‑molecule, earth‑system and medical‑image data via native scientific tokens, delivering competitive benchmark results across life, material, earth and medical domains while enabling seamless cross‑modal inference and open community collaboration.
Artificial intelligence models for scientific discovery are shifting from "engineering stitchers" to "intelligent reasoners".
In recent years, the AI for Science field has produced many high‑performance specialized models for tasks such as protein‑structure prediction, molecular generation and material design, yet breakthroughs often arise at cross‑modal intersections.
For disease research, where gene sequences, protein conformations and pharmacological activity data intertwine, researchers need a unified framework that can simultaneously read sequence language, understand spatial conformation and infer dynamic interactions.
In response, the Shanghai Institute of Scientific Intelligence released the multimodal scientific foundation model “ShenZhen”. With a total parameter count of 11 billion, it is not a simple aggregation of specialist models; it aligns modalities at the representation level within a single architecture and currently supports six scientific data types: DNA, RNA, protein, small molecules, earth‑system data and medical images.
Specialized models such as AlphaFold2 have demonstrated the potential of AI on single scientific tasks, but their strict modality‑specific representation limits generalisation and cross‑domain knowledge transfer, creating data and task islands.
Since 2022, large language models (LLMs) have been used as a universal reasoning backbone by converting scientific data into text tokens (e.g., SMILES strings for molecules). However, flattening inherently structured scientific data into plain text loses systematic information, capping inference accuracy.
Professor Qi Yuan, director of the institute, emphasized at WAIC 2026 that a good model must compress high‑dimensional space without discarding critical details, allowing AI to move from token prediction to discovering unknown laws.
Guided by this view, the team proposes a new scientific‑AI approach: ingest scientific modalities in their native, lossless form, encode them into "scientific tokens" via discipline‑specific tokenizers, and align them in a shared representation space before feeding them to a large language model for reasoning.
The architecture consists of a shared intelligent core plus expert interface modules. Each discipline has a dedicated tokenizer that efficiently compresses high‑dimensional raw data into scientific tokens while preserving order, connectivity, spatial fields or texture details. These tokens are unified in a common space, where the model dynamically invokes appropriate functional modules for cross‑modal understanding, reasoning and generation.
On the output side, modality‑specific decoders translate the shared representation back into structured scientific objects, enabling direct generation of RNA sequences, SMILES strings, global weather fields and medical‑image segmentation masks.
Three practical advantages arise: (1) knowledge sharing and law transfer within a scientific domain (e.g., DNA‑RNA‑protein relationships); (2) reliable cross‑modal inference forming a closed loop for multimodal scientific discovery; (3) a scalable general‑reasoning engine that can be extended.
Comprehensive evaluation on 49 benchmarks across life, material, earth and medical domains shows strong performance. In life‑science tasks (20 benchmarks covering DNA, RNA, protein and sequence relationship judgments), ShenZhen achieved the best result on 9 tasks and ranked in the top two on 17, outperforming larger models such as Intern‑S1‑Pro (≈1 trillion parameters). In material‑science benchmarks (SMolInstruct, 6 small‑molecule property tasks), ShenZhen attained the best or tied‑best result on 4 tasks. For weather forecasting, a rolling 6‑hour interval prediction up to day 10 matched or exceeded operational numerical forecasts on metrics such as 500 hPa geopotential height and 2‑m temperature. In medical‑image segmentation (9 modalities, >100 k samples), ShenZhen achieved an average Dice score of 91.20, the highest among seven competing methods, ranking first on five modalities and second on the remaining four.
A cross‑modal inference case—predicting drug‑target binding affinity—illustrates how ShenZhen replaces a three‑step pipeline (separate encoders, fusion model, glue code) with a single natural‑language interaction: the researcher inputs protein sequence, molecular structure and a textual command, and the model returns the affinity prediction directly.
By open‑sourcing ShenZhen, the institute invites scientists, developers and researchers from diverse disciplines to extend the framework with new scientific modalities, fostering a collaborative ecosystem.
ShenZhen occupies the middle layer of the Xinghe Qizhi scientific‑intelligence platform, which comprises a bottom layer of foundational data, models, tools and skills, the ShenZhen multimodal foundation model, and a top layer featuring the “DaSheng” AI agent that orchestrates end‑to‑end research workflows.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
