S1-Omni Unifies Scientific Multimodal Reasoning
Jiahao Zhao · hf · 2026-07-20
S1-Omni: A Unified Scientific Multimodal Reasoning Model
S1-Omni is a unified multimodal reasoning model designed for AI for Science, aiming to integrate scientific understanding, prediction, and generation within a single model. It maps natural language instructions alongside CIF, SMILES, protein sequences, spectra, and scientific images into a shared representation space, training the model with knowledge of both the natural world and scientific laws.
The authors state that the S1-Omni-Corpus covers 200 scientific tasks, contains millions of reasoning samples, and is evaluated on 60+ scientific benchmarks. Reportedly, it surpasses GPT-5.5 and Gemini-3.1-Pro on most benchmarks and matches or exceeds domain-specific models on several tasks.
Supported tasks include:
- Property prediction
- Spectrum-to-molecule generation
- Protein site and structure prediction
- Scientific image generation and editing
Related event: S1-Omni: A Unified Multimodal Model for Scientific Reasoning(2 posts)→
More from Multimodal
- Pablo Stanley shares a full AI video workflow using ChatGPT, Gemini, Runway and CapCut — jdjohnson · 2026-07-21
- Meta AI text input now lets users interleave images with text — ezyang · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- Same prompt, Seedance 2 and Grok are compared on cinematic transformation output — LudovicCreator · 2026-07-21
- CG Chefs Showcases Retro Anime Style AI Video Generation — nicolascraske · 2026-07-21
- Night-party video demo uses Seedance 2.0, timecode prompts and 4K upscaling — gen_ericai · 2026-07-21