Intern-S2-397B Multimodal Preview Released
burny_tech · x · 2026-07-17
Intern-S2-Preview-397B 是一个面向科学智能和长程智能体的多模态基础模型。帖子强调它在通用推理、科学问题求解和 agent 能力上有明显提升。\n\n核心方法包括:\n- 新的预训练范式,保留文本-视觉对应关系,并增强空间与视觉推理,同时提升数据效率。\n- 通过在 20+ 个领域上联合训练多样化科学 RL 任务,提升通用推理和科学任务表现。\n- 通过把多个 agent 框架连接到大规模 sandbox 环境做黑盒 agentic RL,增强长程任务的泛化与能力上限。\n\n帖中还给出了模型在 Hugging Face 和 ModelScope 上的链接。
Related event: Intern-S2-Preview-397B Appears on Hugging Face(2 posts)→
More from Multimodal
- Video Models Cut Ad Production Costs by 90-99%: Runway Enterprise Data — c_valenzuelab · 2026-07-21
- Pablo Stanley shares a full AI video workflow using ChatGPT, Gemini, Runway and CapCut — jdjohnson · 2026-07-21
- Meta AI text input now lets users interleave images with text — ezyang · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- Same prompt, Seedance 2 and Grok are compared on cinematic transformation output — LudovicCreator · 2026-07-21
- CG Chefs Showcases Retro Anime Style AI Video Generation — nicolascraske · 2026-07-21