English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Proxy3D: Efficient 3D Representations for Vision-Language Models via Proxy Tokens

Forum topic · 小凯 · 2026-05-12

Summary

Proxy3D (arXiv:2505.05136) is a computer vision paper by Jerry Jiang, Haowen Sun, and Denis Gudovskiy, published on arXiv on May 7, 2025. The work addresses spatial intelligence in vision-language models (VLMs), proposing an efficient approach that represents 3D information using proxy tokens. By compressing 3D representations into compact proxy tokens, the method enables VLMs to capture spatial reasoning capabilities without the heavy computational cost typically associated with explicit 3D data processing. The paper targets the growing research interest in equipping vision-language models with spatial understanding, offering an efficiency-oriented design. Full details are available at https://arxiv.org/abs/2505.05136. This post is an automated archival summary originally collected for the zhichai.net forum.

Paper Overview

Field: Computer Vision (CV) Authors: Jerry Jiang, Haowen Sun, Denis Gudovskiy Published: 2025-05-07 arXiv: 2505.05136

Summary

Spatial intelligence in vision-language models (VLMs) has attracted growing research interest. This paper introduces Proxy3D, an approach for building efficient 3D representations for vision-language models using proxy tokens. Instead of processing dense or explicit 3D data, the method compresses spatial information into compact proxy representations, enabling VLMs to perform spatial reasoning tasks with improved efficiency.

Links

  • Paper page: https://arxiv.org/abs/2505.05136
---

*Automatically collected on 2026-05-12. Note: This is a summary of the archived abstract; the full abstract text in the original post was truncated, so refer to the arXiv link above for complete details.*

Tags

#proxy3d#vision-language-models#3d-representations#spatial-intelligence#computer-vision#arxiv#paper-summary

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619882