English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation

Forum topic · 小凯 · 2026-07-09

Summary

Lift3D-VLA (arXiv:2507.06837) is a unified Vision-Language-Action (VLA) framework for robotic manipulation that adds explicit 3D point cloud reasoning and temporally coherent action generation. It introduces an enhanced 2D model-lifting strategy that geometrically aligns 3D points with pretrained 2D positional embeddings, allowing the VLA vision encoder to encode point clouds directly with minimal spatial information loss. A geometry-centric masked autoencoder (GC-MAE) provides dual-objective self-supervised learning: reconstructing current point clouds while predicting their future geometric evolution, so the encoder internalizes 3D structure and physical dynamics. Hierarchical temporal action modeling leverages multiple LLM layers to predict action chunks with temporal consistency. Across 22 simulated tasks and 8 real-world manipulation tasks, Lift3D-VLA outperforms prior best VLA methods by 10.8% average success rate on MetaWorld and 11.1% on RLBench, surpasses the strongest real-world baseline by 4 percentage points, and shows stronger generalization to out-of-distribution perturbations.

Paper Overview

Field: Computer Vision (CV) Authors: Jiaming Liu, Qingpo Wuwu, Nuowei Han Published: 2025-07-09 arXiv: 2507.06837

Abstract (English)

Recently, Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse tasks. However, effective robotic manipulation in physical environments fundamentally requires geometric understanding and spatial reasoning. While some VLA approaches attempt to incorporate 3D information, they are constrained by limited data availability and geometric information loss in current 3D encoding pipelines, and fail to jointly capture 3D geometry and temporally structured actions in dynamic environments. To address these limitations, the authors introduce Lift3D-VLA, a unified VLA framework that equips models with explicit 3D point cloud reasoning and enables temporally coherent action generation.

Key Contributions

  • Enhanced 2D model-lifting strategy: Building upon the authors' previous work Lift3D, 3D points are geometrically aligned with pretrained 2D positional embeddings, enabling the VLA vision encoder to directly encode point clouds while minimizing spatial information loss.
  • Geometry-Centric Masked Autoencoding (GC-MAE): A dual-objective self-supervised framework that reconstructs the current point cloud while simultaneously predicting its future geometric evolution, allowing the 2D vision encoder to internalize 3D structure and physical dynamics.
  • Hierarchical temporal action modeling: Multiple layers of the LLM collaborate to predict action chunks, achieving temporally consistent action generation.
  • Results

    Evaluated on 22 simulated tasks and 8 real-world manipulation tasks, Lift3D-VLA achieves:

  • +10.8% average success rate over the previous best VLA methods on MetaWorld
  • +11.1% average success rate over previous best VLA methods on RLBench
  • +4 percentage points over the strongest real-world baseline
  • Stronger generalization to out-of-distribution perturbations
  • Links

  • arXiv: https://arxiv.org/abs/2507.06837
---

*Auto-collected on 2026-07-09.*

Tags

#robotics#vla#3d-perception#manipulation#self-supervised-learning#arxiv#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346256