English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MCF-Net: Motion-Conditioned Multi-View Fusion for Myocardial Infarction Localization in Echocardiography

Forum topic · 小凯 · 2026-07-18

Summary

Researchers propose MCF-Net, a motion-guided multi-view fusion framework for localizing myocardial infarction (MI) from echocardiography (Echo). While foundation models have improved vision-based Echo analysis, most existing methods operate on single views, and segment-level localization remains unreliable under view-dependent ambiguity, particularly in apical views. MCF-Net fuses myocardial motion cues with foundation model representations: visual features are extracted with EchoPrime, a pretrained Echo foundation model shared across two views. Cardiac motion is modeled with extremely sparse supervision—a single annotated template frame is propagated across videos to initialize point tracking, avoiding dense annotation. Motion-derived segment-aware soft masks provide coarse spatial priors to enhance features of challenging myocardial segments, and a motion-conditioned fusion mechanism integrates motion and vision across views without overriding strong appearance cues. On segment-level MI localization, MCF-Net achieves 72.4% F1 score and 84.9% accuracy, outperforming state-of-the-art motion-only, vision-only, and fusion baselines. Paper: arXiv 2607.15268.

Paper Overview

Field: Computer Vision Authors: Guang Yang, Wentian Xu, Siyu Wang, Betty Raman, Lei Li, Vicente Grau Published: 2026-07-16 arXiv: 2607.15268

Abstract

Myocardial infarction (MI) remains a leading cause of mortality worldwide. Echocardiography (Echo) is a widely available modality for MI assessment, where regional wall motion abnormality is a key indicator. Prior learning-based methods for myocardial motion analysis often use handcrafted descriptors or densely supervised estimation, but the need for extensive annotation limits applicability.

Foundation models have recently improved vision-based Echo analysis; however, most methods operate on single views, and segment-level localization remains unreliable under view-dependent ambiguity, especially in apical views.

Method: MCF-Net

To address these limitations, the authors propose MCF-Net, a novel motion-guided multi-view fusion framework that fuses myocardial motion cues with foundation model representations to localize infarction:

  • Visual features are extracted using EchoPrime, a pretrained Echo foundation model with representations shared across two views.
  • Cardiac motion is modeled with extremely sparse supervision: a single annotated template frame is propagated across videos to initialize point tracking, avoiding the need for dense annotation.
  • Motion-derived segment-aware soft masks provide coarse spatial priors that selectively enhance features of challenging myocardial segments.
  • A motion-conditioned fusion mechanism integrates motion and visual information across views, refining predictions without overriding strong appearance cues.
  • Results

    On segment-level MI localization, MCF-Net achieves:

  • 72.4% F1 score
  • 84.9% accuracy
These results outperform state-of-the-art motion-only, vision-only, and fusion baselines.

---

*Auto-collected on 2026-07-18.*

Tags

#medical-imaging#echocardiography#deep-learning#multi-view-fusion#foundation-models#myocardial-infarction#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178433582