English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ZipDepth: Lightweight Zero-Shot Monocular Depth Estimation for Embedded and Mobile Devices

Forum topic · 小凯 · 2026-07-11

Summary

ZipDepth (arXiv 2507.08183) is a compact monocular depth estimation network by Fabio Tosi, Luca Bartolomei, and Matteo Poggi that brings zero-shot depth inference to resource-constrained platforms. While foundation models achieve robust zero-shot generalization in monocular depth estimation, their computational demands exceed embedded and mobile hardware; existing lightweight models are trained almost exclusively in single-domain self-supervised setups and fail silently under domain shift. ZipDepth bridges this gap by combining an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model, trained on a large multi-domain dataset. With only 6.1M parameters, ZipDepth runs in real time from server GPUs down to power-constrained devices, and achieves the best trade-off between zero-shot accuracy and deployment efficiency among lightweight models across five benchmarks, approaching the accuracy of foundation models 50 times its size.

Overview

  • Field: Computer Vision
  • Authors: Fabio Tosi, Luca Bartolomei, Matteo Poggi
  • arXiv: 2507.08183
  • Key Points

  • Monocular depth estimation has advanced rapidly via foundation models with strong zero-shot generalization, but their compute requirements rule out embedded and mobile platforms.
  • Existing lightweight depth networks were developed almost exclusively within single-domain, self-supervised paradigms and fail silently under domain shift.
  • ZipDepth combines an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model, trained over a large multi-domain dataset.
  • With just 6.1M parameters, ZipDepth runs at real-time rates on everything from server GPUs to power-constrained devices.
  • Across five benchmarks, it achieves the best trade-off between zero-shot accuracy and deployment efficiency among lightweight models, taking a significant step toward the accuracy of foundation models 50x its size.

Original Abstract (excerpt)

> Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond the reach of embedded and mobile platforms. Lightweight alternatives exist, but have been developed almost exclusively within single-domain, self-supervised paradigms, failing silently under domain shift. We present ZipDepth, a compact monocular depth network that bridges this gap by combining an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model over a large multi-domain training set. Comprising just 6.1M parameters, ZipDepth runs at real-time rates from server GPUs to power-constrained devices, achieving the best trade-off between zero-shot accuracy and deployment efficiency...

*Auto-collected on 2026-07-11.*

Tags

#monocular-depth-estimation#zero-shot-learning#knowledge-distillation#lightweight-models#edge-deployment#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346310