English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

The Matching Principle: A Geometric Theory of Loss Functions for Nuisance-Robust Representation Learning

Forum topic · 小凯 · 2026-05-25

Summary

This paper argues that widely treated as separate problems—robustness, domain adaptation, invariance to photometric/occlusion perturbations, temporal robustness, alignment safety, and classical anisotropic regularization—share a common statistical structure. The authors formalize this as the Matching Principle, which states that encoders should be regularized so that the row space of the Jacobian covers the covariance of label-preserving deployment nuisances. CORAL, adversarial training, IRM, data augmentation, metric learning, Jacobian penalties, and alignment-style constraints are reframed as different estimators of this same object. In a linear-Gaussian model the paper proves closed-form optimality (Theorem A), including a cubic-root water-filling solution over the matching subspace; shows range-coverage necessity of quadratic Jacobian penalties (Theorem G); establishes the same range dichotomy at deep global minima; and provides two falsifiability controls plus seven consistency lemmas. A new Trajectory Deviation Index (TDI) serves as an unlabeled probe of embedding sensitivity. Across thirteen pre-registered blocks from classic ML to Qwen2.5-7B, twelve confirm the predicted matching-then-isotropy-then-W-ordering pattern; the sole exception (Office-31) was pre-named as an eigengap failure. At 7B scale, matching-style PMH improves selective honesty and preserves style TDI, whereas standard DPO degrades it.

Paper Overview

Research area: ML Author: Vishal Rajput Posted: 2026-05-25 arXiv: 2505.14491

Summary

Robustness, domain adaptation, invariance to photometric and occlusive perturbations, compositional generalization, temporal robustness, alignment safety, and classical anisotropic regularization are usually treated as independent problems with their own families of methods. This paper argues that most of the shared structure across these areas is fundamentally a statistical problem: estimate the covariance of label-preserving deployment nuisances, then regularize the encoder Jacobian so that the row space of its matrix product covers this covariance (the Matching Principle). CORAL, adversarial training, IRM, data augmentation, metric learning, Jacobian penalties, and alignment-style constraints are all different estimators of this single object rather than independent robustness tricks.

In a linear-Gaussian model, the paper proves:

  • Closed-form optimality (Theorem A), including a cubic-root water-filling solution over the matching subspace.
  • Necessity of range coverage for quadratic Jacobian penalties (Theorem G).
  • The same range dichotomy holds at deep global minima.
  • Two falsifiability controls (Lemma C; Corollary E) and seven consistency lemmas (D1–D7) for estimation under standard identifiability assumptions.
The paper introduces the Trajectory Deviation Index (TDI), an unlabeled probe of embedding sensitivity used when task accuracy or Jacobian Frobenius norms are insufficient.

Across thirteen pre-registered blocks ranging from classic ML to Qwen2.5-7B, the predicted ordering—matching, then isotropy, then W—is tested under geometric and deployment-drift conditions; twelve blocks pass, with the sole exception (Office-31) being a pre-named eigengap failure. At the 7B scale, matching-style PMH improves selective honesty and preserves style TDI, while standard DPO degrades it.

The contribution is naming the deployment-nuisance covariance, stating what regularizers must do, and providing a closed-form, falsifiable theory once that object is identified—rather than claiming universal superiority on every leaderboard.

--- *Auto-collected 2026-05-25*

Tags

#matching-principle#representation-learning#robustness#domain-adaptation#jacobian-regularization#nuisance-covariance#alignment#falsifiability

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620764