Visual Pretraining Enhances Language Intelligence

Yiming Zhang · hf · 2026-07-13

This work challenges the default assumption that language model pretraining must rely solely on plain text, proposing Visual Pretraining as a scalable learning paradigm for foundation models.

The authors argue that visual information in documents and web pages—such as charts, formulas, and layouts—carries substantial knowledge that is lost when converted to plain text. To address this, the paper systematically explores various unsupervised visual pretraining methods that utilize visually-rich documents directly, bypassing text extraction.

Experimental results show that across multiple backbones and benchmarks, visual pretraining using the same corpus consistently outperforms plain text pretraining. The authors conclude that visual pretraining can serve as a more efficient scaling path for language intelligence.

Related event: Scalable Visual Pretraining Boosts Language Intelligence(4 posts)→

Original post →

More from Research

Research channel →