Intern-S2-Preview: Scientific Agentic Foundation Model
Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, Yanhui Duan, Yue Fan, Youqing Fang, Quan Gan, Yuanyuan Gao, Jiaye Ge, Lixin Gu, Yuzhe Gu, Qipeng Guo, Junjun He, Xin Hong, Ming Hu, Zhouqi Hua, Haian Huang, Junhao Huang, Zixian Huang, Minxi Jin, Lingkai Kong, Alexander Lam, Zehao Li, Zonglin Li, Tianhao Liang, Dahua Lin, Junyao Lin, Tianyang Lin, Zhouhan Lin, Jiangning Liu, Jin Liu, Kuikun Liu, Wenran Liu, Yifei Liu, Yuhong Liu, Yuhong Liu, Zhoumianze Liu, Ziyan Liu, Ziyu Liu, Haijun Lv, Han Lv, Chengqi Lyu, Le Ma, Ningsheng Ma, Zerun Ma, Haoyang Peng, Runyu Peng, Jifei Shan, Zixin Shang, Kou Shi, Xiang Shi, Qisheng Su, Xuerui Su, Hao Sun, Xiao Sun, Yanan Sun, Yu Sun, Huanze Tang, Yinghao Tang, Wenhui Tian, Zhongbo Tian, Bingli Wang, Haomin Wang, Jiarui Wang, Jingzhi Wang, Rui Wang, Xiquan Wang, Yi Wang, Zhecan Wang, Ziyi Wang, Zun Wang, Rubin Wei, Lianyi Wu, Wen Wu, Yue Wu, Yuhan Wu, Zhenyu Wu, Zijian Wu, Shuhao Xing, Jun Xu, Xingle Xu, Xuenan Xu, Xiangchao Yan, Ziang Yan, Bowen Yang, Danni Yang, Lin Yang, Zhiqi Yang, Qian Yao, Haochen Ye, Peng Ye, Jinhui Yin, Jiashuo Yu, Dingbo Yuan, Fei Yuan, Yuhang Zang, Bo Zhang, Chao Zhang, Chen Zhang, Hongjie Zhang, Junming Zhang, Wenlong Zhang, Wenwei Zhang, Yiming Zhang, Zhuo Zhang, Ziyang Zhang, Haiteng Zhao, Penghao Zhao, Yibo Zhao, Zhonghan Zhao, Zhihang Zhong, Bowen Zhou, Peiheng Zhou, Xin Zhou, Xinyu Zhou, Yunhua Zhou, Dongsheng Zhu, Yicheng Zou
cs.LG, cs.CL, cs.CV
2026-08-14
Shanghai AI Lab's 397B Intern-S2-Preview combines scientific multimodal pretraining with agentic RL, and a pluggable 4B memory module lifts its Biology-Instructions score from 56.92 to 60.32 without touching the frozen backbone.
Scientific discovery demands more than answering a single question. It requires reading heterogeneous evidence like charts, spectra, and satellite imagery, invoking tools, and sustaining progress across long task horizons. Existing models split into two incomplete camps: general-purpose LLMs are strong at instruction following and reasoning but aren't specialized for scientific modalities, domain protocols, or verifiable tool interaction, while scientific multimodal models handle specialized perception well but are usually evaluated as static question-answering systems rather than long-horizon agents. Shanghai AI Laboratory's Intern-S2-Preview tries to bring both threads together: a foundation model that understands scientific multimodal evidence and can also act as an agent over extended workflows.
Training runs in two stages. Pretraining focuses on scientific multimodal data: visual pretraining learns directly from rendered scientific document pages by predicting visual latents, recovering layout information that plain text extraction throws away; interleaved image-text data is built by parsing pages, cropping visually informative regions, and reassembling text and images in layout-aware sequences; a large-scale image-retrieval pipeline recalls and reranks high-quality scientific images for multimodal training.
Post-training is a unified pipeline. Supervised fine-tuning first establishes instruction-following and tool-use behavior, then scalable multi-task reinforcement learning optimizes reasoning depth, correctness, scientific generation quality, and response efficiency under verifiable objectives, using Group-level Entropy-Controlled Policy Optimization (GEPO) to balance exploration across task groups with different entropy regimes. For long-horizon agentic tasks, the team built a black- and white-box agentic RL framework based on a harness times task abstraction that decouples agent runtimes from executable task distributions, letting different tool-using agents and tasks share one rollout, verification, and training protocol; tasks are drawn from public coding and terminal benchmarks plus a self-evolving task-synthesis system built on community skills. On-policy distillation finally merges the separately trained reasoning and agentic expert policies into the unified model. Several systems techniques keep this pipeline stable and efficient underneath: partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks.
Architecturally, two extensions stand out. The 397B backbone gets a dedicated forecasting branch that extends its time-series modeling from long-sequence understanding to numerical forecasting. Separately, a Memory Decoder is trained to attach external parametric memory to the frozen 397B backbone, adding domain knowledge without touching the backbone's weights.
On scientific benchmarks, Intern-S2-Preview-397B outperforms the compared open- and closed-source models on Biology-Instructions (56.92), Mol-Instructions (52.37), and SciReasoner (63.97), and tops the internal MP20 and ProteinBinder-9 sets; it's the best open-source model on MolecularIQ, TOMG-Bench, XLRS-Bench, and MicroVQA. On general benchmarks it leads open-source models on MMLU-Pro (89.75), SimpleQA-Verified (69.90), MMMU-Pro (80.46), and ChartQAPro (69.65), though it still trails closed flagships like Gemini-3.1-Pro on some of these (75.60 on SimpleQA-Verified). On agentic tasks (SkillsBench, TerminalBench 2.1, SWE-Bench-Pro, SWE-Bench-Multilingual, WildClawBench) it generally beats DeepSeek-V4-pro and Qwen3.5-397B but ranks behind GLM-5.2.
On time-series understanding, it clearly beats general-purpose text and vision-language LLMs on SciTS, pushing the PHU01 task's F1 from 36.8, achieved by the larger trillion-parameter Intern-S1-Pro, to 66.9, despite having less than half the parameters. On forecasting, the dedicated numerical branch outperforms specialized time-series baselines on several SciTS subtasks, hits 99% accuracy on predicting the required horizon length, and achieves a competitive zero-shot MASE of 0.785 on the general-purpose GIFT-Eval benchmark. In the Memory Decoder study, a separately trained 4B-parameter biology memory (Intern-MemDec-4B) attached to the frozen 397B backbone raises the Biology-Instructions average across 21 tasks from 56.92 to 60.32, while leaving performance on general knowledge, reasoning, and multimodal benchmarks essentially unchanged, meaning the add-on memory doesn't cost the backbone its general capability.
The interesting part here isn't any single leaderboard win, it's that the team turned scientific multimodal perception, long-horizon agentic capability, and domain extensibility into one reproducible training pipeline. The Memory Decoder path is worth watching in particular: if attaching a few-billion-parameter module to a frozen backbone can meaningfully lift a scientific subdomain without hurting general capability, that's a much cheaper alternative to fully fine-tuning a 397B model every time a new domain shows up, a realistic option for resource-constrained labs or vertical teams. The harness-task decoupling in the agentic RL framework is also a reusable idea: separating the agent runtime from the task distribution makes it easier to keep adding new tasks without redesigning the training protocol each time.
The authors describe Intern-S2-Preview as still a preview system, with future work needed on reliability over longer scientific workflows, expanding domain memories and task environments, strengthening verifiers, and deepening integration with specialized scientific tools. What the paper doesn't emphasize: on general-purpose benchmarks it still trails top closed models like GPT-5.5 and Gemini-3.1-Pro by a clear margin, and on agentic tasks it ranks behind GLM-5.2, so the scientific-agent pitch rests more on domain-specific science and time-series benchmarks than on beating top general models outright. The Memory Decoder validation is also limited to a single domain, biology; whether the same gains generalize to other scientific fields like materials science or earth science isn't demonstrated here.