New Visual Pretraining Paradigm Beats Text with 1/4 Tokens
量子位 · wechat · 2026-07-31
USTC and the Shanghai Artificial Intelligence Laboratory proposed a novel Visual Pretraining (VP) paradigm. Research shows that in real scientific corpora, images carry visual logic that cannot be losslessly translated into text.
- Core Breakthrough: By 'predicting the next visual unit' in a continuous visual representation space, models can learn directly from interleaved image-text documents. Compared to traditional text pretraining, VP achieves better results using only about a quarter of the tokens.
- Capability Boost: VP not only outperforms text pretraining on pure-text reasoning benchmarks like MMLU-Pro but also shows a clear Scaling law trend. Furthermore, it effectively promotes cross-modal alignment without the need for explicit image-text pair supervision.
- Application: This technology has been applied to the pretraining of the Intern-S2-Preview series of multimodal large models. Completed on domestic Ascend computing platforms, it significantly enhances the model's scientific reasoning and multimodal understanding capabilities.
More from Research
- Deep20Bench tests LLM strategy via 'Twenty Questions': Opus 5 and Kimi K3 lead the pack — wauwau0977 · 2026-07-31
- GitHub Hit 'Data Structures in Practice': A Hardware-Aware Guide for System Engineers — tom_doerr · 2026-07-31
- Paper Reveals Deep Research Vulnerability: Misleading Info Triggers False Conclusions — Pengyu Zhu · 2026-07-31
- Study: Simple Image Transformations Easily Bypass Commercial AI Content Moderation — chaumian · 2026-07-31
- Pangram-4 Tech Report: Training a SOTA AI-Text Detector with Repeat2 Trick — RexDouglass · 2026-07-31
- Kimi K3 Architecture Explained: Building a 2.8T Parameter Open Model via 'Active Forgetting' — AccBalanced · 2026-07-31