New Visual Pretraining Paradigm Beats Text with 1/4 Tokens

量子位 · wechat · 2026-07-31

USTC and the Shanghai Artificial Intelligence Laboratory proposed a novel Visual Pretraining (VP) paradigm. Research shows that in real scientific corpora, images carry visual logic that cannot be losslessly translated into text.

Original post →

More from Research

Research channel →