Paper Overview
Field: Machine Learning
Authors: ML Nissen Gonzalez, Melwina Albuquerque, Laurence Wroe, Jacob Meyer Cohen, Logan Riggs Smith, Thomas Dooms
Published: 2026-05-14
arXiv: 2605.15183
Abstract
Mechanistic interpretability aims to break models into meaningful parts; verifying that two such parts implement the same computation is a prerequisite. Existing similarity measures evaluate either empirical behaviour, leaving them blind to out-of-distribution mechanisms, or basis-dependent parameters, meaning they disregard weight-space symmetries.
To address these issues for the class of tensor-based models, the authors introduce a weight-based metric, tensor similarity, that is invariant to such symmetries. This metric captures global functional equivalence and accounts for cross-layer mechanisms using an efficient recursive algorithm.
Empirically, tensor similarity tracks functional training dynamics, such as grokking and backdoor insertion, with higher fidelity than existing metrics.
Key Points
- Verifying computational equivalence between network components is a prerequisite for mechanistic interpretability
- Behavior-based metrics miss out-of-distribution mechanisms; parameter-based metrics are sensitive to weight-space symmetries
- Tensor similarity is a weight-based metric invariant to these symmetries
- It captures global functional equivalence including cross-layer mechanisms via an efficient recursive algorithm
- Empirically tracks grokking and backdoor insertion dynamics more faithfully than prior metrics
*Source: arXiv:2605.15183*