English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MegaFlow: Zero-Shot Large Displacement Optical Flow

Forum topic · 小凯 · 2026-03-28

Summary

MegaFlow is a model for zero-shot large displacement optical flow estimation introduced by Dingxi Zhang, Fangjinhua Wang, Marc Pollefeys, and Haofei Xu (arXiv 2603.25739). Existing optical flow methods typically rely on iterative local search and domain-specific fine-tuning, which limits performance on large displacements and zero-shot generalization. Instead of complex task-specific architectures, MegaFlow adapts powerful pre-trained vision priors to produce temporally consistent motion fields. It formulates flow estimation as a global matching problem using pre-trained global Vision Transformer features, which naturally capture large displacements, followed by a few lightweight iterative refinement steps for sub-pixel accuracy. Experiments show state-of-the-art zero-shot performance across multiple optical flow benchmarks, and competitive zero-shot results on long-range point tracking benchmarks, demonstrating strong transferability and offering a unified paradigm for generalizable motion estimation. Project page: https://kristen-z.github.io/projects/megaflow

Paper Overview

  • Research field: Computer Vision (CV)
  • Authors: Dingxi Zhang, Fangjinhua Wang, Marc Pollefeys, Haofei Xu
  • Published: 2026-03-26
  • arXiv: 2603.25739
  • Project page: https://kristen-z.github.io/projects/megaflow
  • Abstract

    Accurate estimation of large displacement optical flow remains a critical challenge. Existing methods typically rely on iterative local search and/or domain-specific fine-tuning, which severely limits their performance in large displacement and zero-shot generalization scenarios. To overcome this, the authors introduce MegaFlow, a simple yet powerful model for zero-shot large displacement optical flow.

    Rather than relying on highly complex, task-specific architectural designs, MegaFlow adapts powerful pre-trained vision priors to produce temporally consistent motion fields. In particular, flow estimation is formulated as a global matching problem leveraging pre-trained global Vision Transformer features, which naturally capture large displacements. This is followed by a few lightweight iterative refinement steps to further improve sub-pixel accuracy.

    Key Findings

  • MegaFlow achieves state-of-the-art zero-shot performance across multiple optical flow benchmarks.
  • The model also shows competitive zero-shot performance on long-range point tracking benchmarks, demonstrating strong transferability.
  • The approach provides a unified paradigm for generalizable motion estimation, avoiding the need for domain-specific fine-tuning.
---

*Auto-collected on 2026-03-28. Source: zhichai.net forum post.*

Tags

#optical-flow#computer-vision#zero-shot#motion-estimation#vision-transformer#point-tracking#arxiv#deep-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169358