English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

A Scale-Adaptive Framework for Joint Spatiotemporal Super-Resolution of Precipitation

Forum topic · 小凯 · 2026-04-27

Summary

A new paper on arXiv (2604.21903) by Max Defez, Filippo Quarenghi, Mathieu Vrac, Stephan Mandt, and Tom Beucler introduces a scale-adaptive framework for joint spatiotemporal super-resolution (SR) of climate data, targeting limitations of existing video SR models that are typically tuned to a single pair of spatial and temporal upsampling factors. The method decomposes SR into a deterministic prediction with attention-based conditional means plus a residual conditional diffusion model with optional mass conservation (ensuring precipitation totals match between input and output). Scale adaptivity is achieved by rescaling three factor-dependent hyperparameters: the diffusion noise schedule amplitude beta (larger for bigger SR factors to increase diversity), the temporal context length L (maintaining comparable attention range across frame rates), and a tapering function for the mass-conservation term to limit extreme amplification at large factors. The key hypothesis is that larger SR factors increase underdetermination—requiring more context and residual uncertainty—rather than changing the conditional mean structure. Demonstrations on Comephore reanalysis precipitation over France show a single reusable architecture spanning spatial factors from 1 to 25 and temporal factors from 1 to 6.

Overview

  • Field: Machine Learning
  • Authors: Max Defez, Filippo Quarenghi, Mathieu Vrac, Stephan Mandt, Tom Beucler
  • Published: 2026-04-23
  • arXiv: 2604.21903
  • Abstract (translated from the Chinese forum summary)

    Deep learning video super-resolution (SR) has advanced rapidly, but climate applications usually perform SR along only the spatial or temporal dimension. Existing joint spatiotemporal SR models are typically designed for a single pair of SR factors (the spatial and temporal upsampling ratios between the low-resolution and high-resolution sequences), which limits transfer across spatial resolutions and temporal frame rates.

    The authors propose a scale-adaptive framework that decomposes spatiotemporal SR into:

    1. A deterministic prediction of conditional means using attention mechanisms. 2. A residual conditional diffusion model with an optional mass-conservation transform (ensuring the precipitation total is identical in input and output) to preserve aggregate amounts.

    This decomposition lets the same architecture be reused across different SR factor pairs. The framework's core hypothesis is that larger SR factors mainly increase underdetermination—requiring more context and residual uncertainty—rather than changing the structure of the conditional mean. Scale adaptivity is achieved by rescaling three factor-dependent hyperparameters:

  • Diffusion noise schedule amplitude beta: larger for bigger SR factors to increase output diversity.
  • Temporal context length L: set to keep the attention range comparable across frame rates.
  • Mass-conservation function f: tapered to limit extreme amplification at large factors.

Results

Demonstrations on Comephore reanalysis precipitation over France show that a single architecture spans spatial SR factors from 1 to 25 and temporal factors from 1 to 6, yielding a reusable architecture and a tuning scheme for joint spatiotemporal super-resolution across scales.

Link

Paper: https://arxiv.org/abs/2604.21903

--- *Auto-collected on 2026-04-27.*

Tags

#machine-learning#super-resolution#diffusion-models#climate#precipitation#spatiotemporal#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618804