SIIA Releases Science Multimodal Model Shenzhen

机器之心 · wechat · 2026-07-19

Synced introduced **"Shenzhen"**, a scientific multimodal foundation model released and open-sourced by the Shanghai Institute for Advanced Study (SIIA). Built on Qwen3-VL-8B as a shared backbone, it features modality-specific encoders/decoders for six types of scientific data: DNA, RNA, proteins, small molecules, Earth systems, and medical imaging, enabling understanding, reasoning, and generation within a unified representation space. The article emphasizes that rather than forcefully converting scientific data into text tokens, the model ingests data in a "native, lossless" manner and maps it to scientific tokens. This preserves critical information like sequence order, molecular structures, spatial fields, and image textures. Outputs retain their scientific object forms, such as RNA sequences, SMILES, meteorological fields, and segmentation results, allowing direct integration into downstream design, simulation, and validation workflows. Evaluated across 49 benchmarks, the model ranked first in 9 out of 20 life science tasks and reached the top two in 17. It achieved leading results in multiple small molecule tasks and matched or exceeded operational numerical weather forecasts by day 10 globally. For medical image segmentation, it achieved an average Dice score of 91.20, the highest among evaluated methods. The article also highlights its synergy with the "Galaxy" platform and the "Dasheng" scientific research agent.

Original post →

More from Apps

Apps channel →