Paper Overview
- Research area: CV/AI
- arXiv: 2607.09657
- Authors: Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen
- Published: 2026-07-10
Abstract
The advancement of large foundation models has been driven primarily by pretraining on massive text corpora. However, much knowledge is conveyed through visual representations—charts, typeset equations, and page layouts carry rich information that text alone cannot fully capture. Current methods convert visually rich sources into plain text, discarding these cues.
This paper challenges the default assumption that language models must be trained purely on text, demonstrating that visual pretraining is a scalable learner for foundation model intelligence. Through a systematic study, the authors show that visual pretraining consistently outperforms text-only pretraining on the same corpus, offering an efficient path toward scalable language intelligence.
*Auto-collected on 2026-07-14.*