English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VLA vs VLM Deep Comparison and a Survey of Gemini/Gemma Multimodal Architectures

Forum topic · 小凯 · 2026-06-15

Summary

This forum post from zhichai.net presents two in-depth research reports. Part one compares Vision-Language Models (VLMs) and Vision-Language-Action models (VLAs): VLMs output text while VLAs output robot actions. It covers architectural lineages (CLIP, LLaVA, RT-2, OpenVLA, NVIDIA GR00T N1), training data asymmetry (web-scale image-text pairs vs. scarce robot demonstration data), deployment constraints (VLA requires ~10ms control latency), safety implications, hybrid VLM+VLA system designs, and evaluation challenges beyond task success rate. Part two investigates whether Google's Gemini and Gemma are natively multimodal. Based on official technical reports, it concludes Gemini (arXiv:2312.11805) is natively multimodal, Gemma 3 (arXiv:2503.19786) uses a bolted-on frozen SigLIP encoder similar to LLaVA, while Gemma 4 12B shifts to an encoder-free architecture with native audio input. It also frames VLM architecture evolution into three eras and discusses trade-offs between native multimodal and modular adapter-based approaches.

VLA vs VLM Deep Comparison and Gemini/Gemma Architecture Survey

*Translation and structured summary of a zhichai.net forum research post.*

Part 1: VLA vs VLM In-Depth Comparison

Core distinction

  • VLM (Vision-Language Model): takes an image plus text and produces *text* — captions, answers, descriptions.
  • VLA (Vision-Language-Action model): takes images plus instructions and outputs *actions* — joint angles and control commands. Asked to "fetch the red box," a VLM describes where the box is; a VLA directly drives a robot arm.
  • The fundamental line: VLMs output text; VLAs output actions. They are not competitors but a progression — VLM is the foundation, VLA extends it toward embodied intelligence.

    Architectures

    VLM lineage: Most modern VLMs follow a common recipe — a pretrained LLM backbone + a ViT visual encoder + projection layers mapping image features into the LLM's token space. CLIP (2021) pioneered contrastive alignment; Flamingo and BLIP moved toward generation; LLaVA (2023) established the "ViT + projection + LLM" standard, followed by Qwen2-VL and LLaMA 3.2 Vision. Newer efforts like Emu3 explore *native multimodality*, unifying modalities at the token level.

    VLA lineage (built on VLM backbones with modified output heads):

  • End-to-end (RT-1, RT-2, OpenVLA): image → motor commands directly; simple but limited generalization across robots.
  • Dual-system (NVIDIA GR00T N1): a fast reactive "System 1" (~10ms latency) plus a slow reasoning/planning "System 2."
  • Hierarchical (CogACT, NaVILA): LLM plans subgoals; a dedicated low-level controller executes. Modular but latency accumulates.
  • Self-correcting (SC-VLA): fast path by default; LLM activated for diagnosis and recovery on failure.
  • Training data: rich vs. poor

  • VLMs train on internet-scale image-text data (LAION, COCO, Visual Genome — hundreds of millions of pairs) at low cost.
  • VLAs need synchronized robot demonstration data: camera frames, joint angles, gripper states, and trajectory success labels — collected via VR teleoperation, reinforcement learning, or simulation (Sim2Real). Open X-Embodiment is large but far smaller than LAION, limiting VLA generalization.
  • Notably, VLAs place *higher* demands on the visual encoder than VLMs: action generation (e.g., grasp pose prediction) requires far greater spatial precision, and studies suggest adding control-oriented supervision to the visual encoder yields more benefit than tuning the language module.
  • Training methodology

  • VLMs: pretrain on image-text pairs, fine-tune on VQA/captioning tasks (often with LoRA).
  • VLAs: initialize from a pretrained VLM, extend the action vocabulary, and train with regression (or discretized classification) losses over the action space, plus Sim2Real transfer, curriculum learning, and multi-task training.
  • Important caveat: VLM benchmark performance does not predict VLA performance — a high VQA score does not guarantee good robot control, and vice versa.
  • Deployment: latency is a hard constraint

  • VLM outputs are text tokens; seconds of latency are usually acceptable.
  • VLA outputs drive control loops at ~10ms periods; 100ms inference makes stabilization extremely difficult, and latency jitter can make humanoid robots fall. Acceleration research (distillation, quantization, early exit, lightweight architectures) is active.
  • Failure costs differ: a wrong VLM answer is a bad description; a wrong VLA output can damage property or injure people. VLA deployments require safety validation layers, action bounds, and recovery mechanisms.
  • Application scenarios

  • VLM alone: no physical control needed — warehouse item recognition, offline log analysis, semantic scene annotation.
  • VLA: end-to-end control — manipulation, humanoid locomotion, end-to-end autonomous driving, drone navigation.
  • Hybrid (most common in practice): VLM as the "brain" (instruction parsing, task decomposition) + VLA/controllers as the "cerebellum" (motion generation). Modular and easier to debug and verify.
  • Evaluation

    VLM evaluation is mature (VQA accuracy, BLEU/ROUGE, grounding IoU, hallucination benchmarks). VLA evaluation must additionally cover robustness to environmental changes, execution efficiency, error recovery, and safety — requiring real robots or high-fidelity simulators at much higher cost.

    Current bottlenecks

  • VLM: hallucination, alignment, fairness; mitigations include RLHF/DPO, multimodal chain-of-thought, and external fact-checking tools.
  • VLA: (1) data scarcity, (2) Sim2Real gap, (3) safety verification, (4) the latency-vs-performance trade-off.
  • Notable models

  • VLM: CLIP (2021), LLaVA series, Qwen2-VL/Qwen3, Emu3 (2024).
  • VLA: RT-1/RT-2 (2022–2023), OpenVLA (2024), GR00T N1 (2025), Pi-0 (2024), plus domain-specific VLAs (CoVLA, OpenDriveVLA for driving; medical and agricultural variants).
  • > One-line summary: VLMs understand the world and speak; VLAs understand the world and act. VLM output is text; VLA output is action.

    Part 2: Gemini and Gemma Architecture Survey

    What "native multimodality" means

    A natively multimodal model is jointly trained from the start on interleaved multimodal data, rather than assembling separately pretrained components. Key criteria vs. LLM+ViT bolt-on designs:

    | Dimension | Bolt-on (non-native) | Native multimodal | |---|---|---| | Training | Train ViT and LLM separately, then train a projector | Joint pretraining on interleaved text/image/audio/video | | Encoder | Independent (often frozen) encoder + projection | No frozen encoder; unified Transformer or lightweight embedding | | Fusion depth | Shallow, only at projection | Cross-modal attention in all Transformer layers | | Modality status | Visual features "translated" for the LLM | All modality tokens equal |

    Gemini: natively multimodal

    Google DeepMind's Gemini 1.0 technical report states it trains natively multimodal models over interleaved text, images, audio, and video "from the ground up, not by bolting a frozen vision encoder onto a text decoder." Technical analyses describe:

  • Tokenization: BPE for text; ViT-style patch projection for images; spectrogram/VQ-VAE tokenization for audio; video as spatiotemporal patches with temporal subsampling.
  • Fusion: all tokens enter a single Transformer decoder; cross-modal attention in every layer; arbitrary interleaving.
  • Training: unified next-token autoregressive objective over a flattened multimodal sequence.
  • MoE (Gemini 3.0+): fine-grained routing and modality-aware experts.
  • Caveat: Google has not fully disclosed the architecture; the "natively multimodal" claim is confirmed in official materials, but details come from third-party analyses.

    Gemma 3: bolt-on (not native)

    Per the Gemma 3 technical report (arXiv:2503.19786): a separate frozen 400M SigLIP ViT encoder (images resized to 896×896) shared across the 4B/12B/27B models, producing 256 soft tokens concatenated with text tokens — highly similar to LLaVA. The 1B model has no vision support.

    Gemma 4 12B: encoder-free native multimodality

    Announced 2026-06-03, Gemma 4 12B abandons separate encoders:

  • Vision: a 35M-parameter lightweight embedding module (single matrix multiplication, position embeddings, normalization) replaces the 27-layer ViT; raw 48×48 pixel patches project directly into the LLM's hidden dimension.
  • Audio: the 12-layer Conformer is removed; raw 16kHz audio is framed at 40ms and linearly projected — the first native audio input in a mid-sized Gemma.
  • Fusion: vision, audio, and text share all weights in a unified decoder-only Transformer.
  • Runs in as little as 16GB of VRAM; Apache 2.0 licensed.
  • Three eras of VLM architecture

  • Era 1 (2021–2022): dual towers with learnable bridges (Q-Former) — CLIP, BLIP, Flamingo.
  • Era 2 (2023–2025): pretrained LLM as backbone, vision as a pluggable adapter (MLP/Resampler) — LLaVA, Qwen2.5-VL, early GPT-4V.
  • Era 3 (2025–2026): no bridge; single tokenizer/embedding space; single Transformer trained end-to-end.
  • Era 3a (native input → text output): Gemini 3, Gemma 4, Qwen3.5, GPT-5.4, Claude Opus 4.6.
  • Era 3b (omni-modal unified input/output, with image/audio decoder heads): BAGEL, Qwen3.5-Omni, Emu3.5, Janus-Pro, Ernie 5.0.

Is native multimodality strictly better?

Pros: deep cross-modal fusion in all layers; modality equality; end-to-end optimization. Cons: extremely high training cost (cannot reuse pretrained unimodal components); components cannot be swapped independently. Bolt-on pros: low training cost, independent component upgrades, mature open-source ecosystem — which is why it remains practically valuable.

Conclusions

1. Gemini (1.0+) is genuinely natively multimodal. 2. Gemma 3 is a bolt-on LLM+ViT design, not native. 3. Gemma 4 12B marks a fundamental shift to encoder-free native multimodality. 4. 2026 flagship models have broadly converged on Era 3a native multimodal input architectures.

References

1. *Gemini: A Family of Highly Capable Multimodal Models*, arXiv:2312.11805, Google DeepMind, 2023 2. *Gemma 3 Technical Report*, arXiv:2503.19786, Google DeepMind, 2025 3. Gemma 4 12B Announcement, Google Official Blog, 2026-06-03 4. *A Survey of State of the Art Large Vision Language Models*, arXiv:2501.02189, CVPR 2025 Workshop 5. Vision-Language Models Overview: https://github.com/zli12321/Vision-Language-Models-Overview

Tags

#vla#vlm#multimodal#gemini#gemma#embodied-ai#robotics#model-architecture

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981355