English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Scaling Native Multimodal Pre-Training From Scratch: Tencent & CUHK Uncover the Optimal Recipe for a Bilingual AI Brain

Forum topic · ✨步子哥 · 2026-07-27

Summary

A paper from The Chinese University of Hong Kong and Tencent's LLM Department, 'Scaling Native Multimodal Pre-Training From Scratch' (arXiv:2607.22043), presents the first systematic scaling laws for training multimodal models natively from scratch, without a CLIP- or SigLIP-style vision encoder. Using MoE IsoFLOP experiments, the authors show that language and multimodal objectives follow distinct scaling laws: the compute-optimal allocation for language is nearly insensitive to data composition, while the multimodal allocation depends strongly on the proportion of multimodal data. Because both objectives share one parameter set, the paper characterizes the trade-off via a Pareto frontier over model size, token count, and multimodal data ratio. Experiments trained on 250B text tokens plus 75B multimodal tokens also reveal a surprising positive transfer: native multimodal training improves purely text-based spatial reasoning rather than degrading language ability. The paper includes honest limitations: validation relies only on smoothed training loss, no code is released, and conclusions are specific to MoE architectures.

Scaling Native Multimodal Pre-Training From Scratch: Tencent and CUHK Find the Optimal Recipe for a "Bilingual Brain"

> Paper: Scaling Native Multimodal Pre-Training From Scratch > Authors: Haoyuan Wu, Aoqi Wu, Hai Wang, Jiajia Wu, Jinxiang Ou, Bei Yu > Institutions: The Chinese University of Hong Kong, Tencent LLM Department > arXiv: 2607.22043 (July 24, 2025)

A Counterintuitive Opening

Imagine raising someone fluent in both Chinese and English. Option A: raise them in a purely Chinese environment until age 18, then send them to a U.S. university for four years. Option B: raise them from birth in a household mixing both languages.

Intuitively, Option A seems safer—master one language first, then learn the second. But bilingualism research suggests that Option B produces people whose two languages are not a "translation relationship" but rather share a single conceptual system. When they think in English, the Chinese semantic network is activated too, and vice versa.

Large language model (LLM) multimodal training faces the same choice.

The Current Mainstream: Late-Fusion "Translation"

Today's mainstream multimodal models (GPT-4V, LLaVA, Kimi, etc.) follow Option A. They first train a text-only LLM, then attach a vision encoder (CLIP, SigLIP) and use a projection layer to "translate" visual features into a language the LLM understands.

This is efficient, but has a fundamental asymmetry: visual and language representations are learned separately, on different data distributions and with different optimization objectives. The vision encoder learns via contrastive learning; the LLM learns via autoregressive language modeling. They are like two people with different native languages communicating through a translator.

Native Multimodal: "Bilingual" From Birth

Native multimodal pre-training follows Option B. From the very start, the model trains on a mix of text and images, using a single Transformer, one set of parameters, and one optimization objective—letting visual and language representations grow in the same representation space.

The idea is not new, but no one had answered a key engineering question: given a fixed compute budget, how much should go to model parameters, how much to training tokens, and what proportion to multimodal data?

This paper is the first to give a systematic answer.

Three Independent Scaling Laws

The authors ran IsoFLOP experiments with a MoE (Mixture of Experts) architecture: fixing total compute while varying model size and token count to find the optimal configuration at each budget.

The key finding: the language objective and the multimodal objective follow different scaling laws.

  • The language objective's allocation law is nearly insensitive to data composition. No matter how many images are mixed into training data, the optimal model size and token count for language loss barely change. Language learning is "rigid"—it requires a fixed scale of parameters and tokens, and adding images does not reduce that.
  • The multimodal objective's allocation law depends strongly on data composition. The higher the multimodal data proportion, the more the optimal configuration shifts.
  • It is like language learning: the grammar and vocabulary demands of Chinese are fixed and don't shrink because you also study English—but your English progress depends on how much time you spend in an English environment.

    Pareto Frontier: A "Three-Body Problem" of Compute

    The harder challenge: language and multimodal objectives share the same parameters. Under a fixed total compute budget, you cannot optimize both simultaneously—you must trade off.

    The authors characterize this with a Pareto frontier. Given total compute \(C_{total}\), there exists a set of non-dominated configurations: you cannot improve one objective without hurting the other. This frontier is the engineering "recipe table."

    The paper gives concrete configuration guidance—for a given compute budget, what model size, token count, and multimodal data ratio to use. It is the first "construction blueprint" for native multimodal pre-training.

    Unexpected Bonus: Multimodal Training Boosts Text-Only Spatial Reasoning

    The most surprising finding comes from downstream evaluation.

    Native multimodal pre-training not only gives the model multimodal capability, it also improves purely text-based spatial reasoning. A model trained on images + text outperforms a text-only model on spatial reasoning questions given in text alone.

    This challenges the intuition that multimodal training "diverts" capacity from language learning. The opposite is true—visual spatial information appears to help the model build stronger spatial concept representations, which can be activated even on text-only tasks.

    Like a bilingual person who may reason better even in pure Chinese—because certain conceptual structures from English reinforce Chinese expression.

    Experimental Scale

  • Training data: 250B text tokens (web, books, papers) + 75B multimodal tokens (image-text pairs, interleaved image-text documents)
  • Architecture: MoE decoder-only Transformer, auxiliary-loss-free, with images projected directly into continuous vectors via patch embedding
  • No vision encoder: this is the fundamental difference from the late-fusion paradigm—no CLIP, no SigLIP, only a patch embedding layer
  • Engineering Significance

    The paper's value: it turns native multimodal pre-training from "sounds nice but no one knows how" into an engineering problem with a blueprint.

    For teams training multimodal models (Kimi, Qwen-VL, Emu, SenseNova), the scaling laws offer three directly usable conclusions:

    1. Language learning is rigid—don't expect adding images to save compute on language training. 2. Multimodal learning is elastic—the multimodal data ratio directly sets the ceiling of multimodal capability. 3. Native training has positive transfer—multimodal training not only preserves text ability but also improves spatial reasoning.

    Honest Assessment

    Limitations of the paper:

  • Validation uses only training loss. The authors admit that, "lacking reliable multimodal validation metrics," they use smoothed training loss as a proxy. The relationship between training loss and downstream performance is not linear, so this proxy may over- or under-estimate real effectiveness.
  • No open-source code. The paper comes from Tencent's LLM Department, but code is not released; other teams would need to build their own MoE training framework to reproduce it.
  • MoE conclusions may not apply to dense models. MoE routing may affect how multimodal and language objectives share parameters; dense-model scaling laws could differ.

One-Sentence Summary

The scaling laws for native multimodal pre-training show that language learning is rigid, multimodal learning is elastic, and in the shared parameter space, visual information reinforces language-based spatial reasoning—a "bilingual brain" is not addition, it's multiplication.

---

Paper: https://arxiv.org/abs/2607.22043 HTML version: https://arxiv.org/html/2607.22043v1 Open-source code: not yet available

Tags

#multimodal#scaling-laws#native-multimodal-pretraining#moe#tencent#cuhk#llm#spatial-reasoning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503725