VLA vs VLM Deep Comparison and Gemini/Gemma Architecture Survey
*Translation and structured summary of a zhichai.net forum research post.*
Part 1: VLA vs VLM In-Depth Comparison
Core distinction
- VLM (Vision-Language Model): takes an image plus text and produces *text* — captions, answers, descriptions.
- VLA (Vision-Language-Action model): takes images plus instructions and outputs *actions* — joint angles and control commands. Asked to "fetch the red box," a VLM describes where the box is; a VLA directly drives a robot arm.
- End-to-end (RT-1, RT-2, OpenVLA): image → motor commands directly; simple but limited generalization across robots.
- Dual-system (NVIDIA GR00T N1): a fast reactive "System 1" (~10ms latency) plus a slow reasoning/planning "System 2."
- Hierarchical (CogACT, NaVILA): LLM plans subgoals; a dedicated low-level controller executes. Modular but latency accumulates.
- Self-correcting (SC-VLA): fast path by default; LLM activated for diagnosis and recovery on failure.
- VLMs train on internet-scale image-text data (LAION, COCO, Visual Genome — hundreds of millions of pairs) at low cost.
- VLAs need synchronized robot demonstration data: camera frames, joint angles, gripper states, and trajectory success labels — collected via VR teleoperation, reinforcement learning, or simulation (Sim2Real). Open X-Embodiment is large but far smaller than LAION, limiting VLA generalization.
- Notably, VLAs place *higher* demands on the visual encoder than VLMs: action generation (e.g., grasp pose prediction) requires far greater spatial precision, and studies suggest adding control-oriented supervision to the visual encoder yields more benefit than tuning the language module.
- VLMs: pretrain on image-text pairs, fine-tune on VQA/captioning tasks (often with LoRA).
- VLAs: initialize from a pretrained VLM, extend the action vocabulary, and train with regression (or discretized classification) losses over the action space, plus Sim2Real transfer, curriculum learning, and multi-task training.
- Important caveat: VLM benchmark performance does not predict VLA performance — a high VQA score does not guarantee good robot control, and vice versa.
- VLM outputs are text tokens; seconds of latency are usually acceptable.
- VLA outputs drive control loops at ~10ms periods; 100ms inference makes stabilization extremely difficult, and latency jitter can make humanoid robots fall. Acceleration research (distillation, quantization, early exit, lightweight architectures) is active.
- Failure costs differ: a wrong VLM answer is a bad description; a wrong VLA output can damage property or injure people. VLA deployments require safety validation layers, action bounds, and recovery mechanisms.
- VLM alone: no physical control needed — warehouse item recognition, offline log analysis, semantic scene annotation.
- VLA: end-to-end control — manipulation, humanoid locomotion, end-to-end autonomous driving, drone navigation.
- Hybrid (most common in practice): VLM as the "brain" (instruction parsing, task decomposition) + VLA/controllers as the "cerebellum" (motion generation). Modular and easier to debug and verify.
- VLM: hallucination, alignment, fairness; mitigations include RLHF/DPO, multimodal chain-of-thought, and external fact-checking tools.
- VLA: (1) data scarcity, (2) Sim2Real gap, (3) safety verification, (4) the latency-vs-performance trade-off.
- VLM: CLIP (2021), LLaVA series, Qwen2-VL/Qwen3, Emu3 (2024).
- VLA: RT-1/RT-2 (2022–2023), OpenVLA (2024), GR00T N1 (2025), Pi-0 (2024), plus domain-specific VLAs (CoVLA, OpenDriveVLA for driving; medical and agricultural variants).
- Tokenization: BPE for text; ViT-style patch projection for images; spectrogram/VQ-VAE tokenization for audio; video as spatiotemporal patches with temporal subsampling.
- Fusion: all tokens enter a single Transformer decoder; cross-modal attention in every layer; arbitrary interleaving.
- Training: unified next-token autoregressive objective over a flattened multimodal sequence.
- MoE (Gemini 3.0+): fine-grained routing and modality-aware experts.
- Vision: a 35M-parameter lightweight embedding module (single matrix multiplication, position embeddings, normalization) replaces the 27-layer ViT; raw 48×48 pixel patches project directly into the LLM's hidden dimension.
- Audio: the 12-layer Conformer is removed; raw 16kHz audio is framed at 40ms and linearly projected — the first native audio input in a mid-sized Gemma.
- Fusion: vision, audio, and text share all weights in a unified decoder-only Transformer.
- Runs in as little as 16GB of VRAM; Apache 2.0 licensed.
- Era 1 (2021–2022): dual towers with learnable bridges (Q-Former) — CLIP, BLIP, Flamingo.
- Era 2 (2023–2025): pretrained LLM as backbone, vision as a pluggable adapter (MLP/Resampler) — LLaVA, Qwen2.5-VL, early GPT-4V.
- Era 3 (2025–2026): no bridge; single tokenizer/embedding space; single Transformer trained end-to-end.
- Era 3a (native input → text output): Gemini 3, Gemma 4, Qwen3.5, GPT-5.4, Claude Opus 4.6.
- Era 3b (omni-modal unified input/output, with image/audio decoder heads): BAGEL, Qwen3.5-Omni, Emu3.5, Janus-Pro, Ernie 5.0.
The fundamental line: VLMs output text; VLAs output actions. They are not competitors but a progression — VLM is the foundation, VLA extends it toward embodied intelligence.
Architectures
VLM lineage: Most modern VLMs follow a common recipe — a pretrained LLM backbone + a ViT visual encoder + projection layers mapping image features into the LLM's token space. CLIP (2021) pioneered contrastive alignment; Flamingo and BLIP moved toward generation; LLaVA (2023) established the "ViT + projection + LLM" standard, followed by Qwen2-VL and LLaMA 3.2 Vision. Newer efforts like Emu3 explore *native multimodality*, unifying modalities at the token level.
VLA lineage (built on VLM backbones with modified output heads):
Training data: rich vs. poor
Training methodology
Deployment: latency is a hard constraint
Application scenarios
Evaluation
VLM evaluation is mature (VQA accuracy, BLEU/ROUGE, grounding IoU, hallucination benchmarks). VLA evaluation must additionally cover robustness to environmental changes, execution efficiency, error recovery, and safety — requiring real robots or high-fidelity simulators at much higher cost.
Current bottlenecks
Notable models
> One-line summary: VLMs understand the world and speak; VLAs understand the world and act. VLM output is text; VLA output is action.
Part 2: Gemini and Gemma Architecture Survey
What "native multimodality" means
A natively multimodal model is jointly trained from the start on interleaved multimodal data, rather than assembling separately pretrained components. Key criteria vs. LLM+ViT bolt-on designs:
| Dimension | Bolt-on (non-native) | Native multimodal | |---|---|---| | Training | Train ViT and LLM separately, then train a projector | Joint pretraining on interleaved text/image/audio/video | | Encoder | Independent (often frozen) encoder + projection | No frozen encoder; unified Transformer or lightweight embedding | | Fusion depth | Shallow, only at projection | Cross-modal attention in all Transformer layers | | Modality status | Visual features "translated" for the LLM | All modality tokens equal |
Gemini: natively multimodal
Google DeepMind's Gemini 1.0 technical report states it trains natively multimodal models over interleaved text, images, audio, and video "from the ground up, not by bolting a frozen vision encoder onto a text decoder." Technical analyses describe:
Caveat: Google has not fully disclosed the architecture; the "natively multimodal" claim is confirmed in official materials, but details come from third-party analyses.
Gemma 3: bolt-on (not native)
Per the Gemma 3 technical report (arXiv:2503.19786): a separate frozen 400M SigLIP ViT encoder (images resized to 896×896) shared across the 4B/12B/27B models, producing 256 soft tokens concatenated with text tokens — highly similar to LLaVA. The 1B model has no vision support.
Gemma 4 12B: encoder-free native multimodality
Announced 2026-06-03, Gemma 4 12B abandons separate encoders:
Three eras of VLM architecture
Is native multimodality strictly better?
Pros: deep cross-modal fusion in all layers; modality equality; end-to-end optimization. Cons: extremely high training cost (cannot reuse pretrained unimodal components); components cannot be swapped independently. Bolt-on pros: low training cost, independent component upgrades, mature open-source ecosystem — which is why it remains practically valuable.
Conclusions
1. Gemini (1.0+) is genuinely natively multimodal. 2. Gemma 3 is a bolt-on LLM+ViT design, not native. 3. Gemma 4 12B marks a fundamental shift to encoder-free native multimodality. 4. 2026 flagship models have broadly converged on Era 3a native multimodal input architectures.
References
1. *Gemini: A Family of Highly Capable Multimodal Models*, arXiv:2312.11805, Google DeepMind, 2023 2. *Gemma 3 Technical Report*, arXiv:2503.19786, Google DeepMind, 2025 3. Gemma 4 12B Announcement, Google Official Blog, 2026-06-03 4. *A Survey of State of the Art Large Vision Language Models*, arXiv:2501.02189, CVPR 2025 Workshop 5. Vision-Language Models Overview: https://github.com/zli12321/Vision-Language-Models-Overview