English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

How Distribution Shift Shapes Pretraining Gains in Neural PDE Surrogates: Insights from Airfoil RANS Fine-Tuning

Forum topic · 小凯 · 2026-09-19

Summary

This paper (arXiv:2609.20814) investigates how distribution shift affects the data-efficiency gains of pretraining neural PDE surrogate models for CFD. The authors pretrain a surrogate on 254,909 RANS solutions from one airfoil family and fine-tune it on a new family under two matched-freestream targets: the same Spalart-Allmaras (SA) turbulence model, and SA augmented with an e^N transition model. At a fine-tuning budget of N=1000, the pretrained model matches from-scratch accuracy achieved with 3.25x more samples for the same-SA target and 2.58x more for the transition target; by N=5000, this ordering reverses (1.56x vs 1.86x). Sampling more distinct airfoils reduces error on both targets at N=1000, but only the same-SA target shows gains exceeding sampling variability (3.3x to 4.0x). The study concludes that pretraining value depends jointly on target data budget, target data coverage, and whether source and target differ in modeled physics.

Paper Overview

Field: Machine Learning Authors: Pochinapeddi Sai Bhargav, Nithin Somasekharan, Rohit Sunil Kanchi, Sicheng He, Shaowu Pan arXiv: 2609.20814

Abstract

Pretraining a neural PDE surrogate can reduce the amount of new CFD data needed when geometry or modeled physics changes. However, it remains unclear how different components of distribution shift affect this benefit.

The authors pretrain a surrogate on 254,909 RANS solutions from one airfoil family and fine-tune it on a new family under two target settings with matched freestream ranges: the same Spalart-Allmaras (SA) modeling, and SA with added e^N transition modeling.

Key Findings

  • At N=1000, the pretrained model matches the accuracy of a model trained from scratch on 3.25x as many samples for the same-SA target, but only 2.58x as many for the transition-modeled target.
  • By N=5000, this ordering reverses (1.56x versus 1.86x), showing pretraining benefits are not monotonic with data budget.
  • At N=1000, sampling more distinct airfoils lowers error on both targets, but only for the same-SA target does the gain exceed the observed inter-sample variability (3.3x to 4.0x).

Conclusion

The value of pretraining depends jointly on three factors: the target data budget, the coverage of the target data, and whether the source and target differ in modeled physics. None of these alone determines the outcome.

--- *Auto-collected on 2026-09-19.*

Tags

#machine-learning#neural-pde-surrogate#cfd#distribution-shift#pretraining#rans#airfoil#transfer-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634979