English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Zero-Shot Depth from Defocus: FOSSA and the ZEDD Benchmark

Forum topic · 小凯 · 2026-03-31

Summary

This paper, arXiv:2503.23737 by Yiming Zuo, Hongyu Wen, and Venkat Subramanian (March 2025), addresses zero-shot generalization in Depth from Defocus (DfD), the task of estimating dense metric depth maps from a focus stack. Unlike prior works that overfit to specific datasets, the authors introduce ZEDD, a new real-world DfD benchmark with 8.3x more scenes and significantly higher-quality images and ground-truth depth maps than previous benchmarks. They also propose FOSSA, a Transformer-based architecture featuring a novel stack attention layer with focus distance embedding that enables efficient information exchange across the focus stack. A new training data pipeline converts existing large-scale RGBD datasets into synthetic focus stacks. Experiments show substantial improvements over baselines, reducing error by up to 55.7%. ZEDD is available at https://zedd.cs.princeton.edu, and code and checkpoints at https://github.com/princeton-vl/FOSSA.

Paper Overview

Field: Computer Vision Authors: Yiming Zuo, Hongyu Wen, Venkat Subramanian Published: 2025-03-30 arXiv: 2503.23737

Abstract

Depth from Defocus (DfD) is the task of estimating a dense metric depth map from a focus stack. Unlike previous works overfitting to a certain dataset, this paper focuses on the challenging and practical setting of zero-shot generalization. The authors first propose a new real-world DfD benchmark, ZEDD, which contains 8.3x more scenes and significantly higher quality images and ground-truth depth maps compared to previous benchmarks. They also design a novel network architecture named FOSSA — a Transformer-based architecture with novel designs tailored to the DfD task. The key contribution is a stack attention layer with a focus distance embedding, allowing efficient information exchange across the focus stack. Finally, they develop a new training data pipeline that allows utilizing existing large-scale RGBD datasets to generate synthetic focus stacks.

Experimental results on ZEDD and other benchmarks show significant improvements over baselines, with error reductions of up to 55.7%.

Resources

  • ZEDD benchmark: https://zedd.cs.princeton.edu
  • Code and checkpoints: https://github.com/princeton-vl/FOSSA
---

*Auto-collected on 2026-03-31*

Tags

#depth-from-defocus#computer-vision#zero-shot-learning#transformer#depth-estimation#benchmark#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169447