English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability

Forum topic · 小凯 · 2026-05-15

Summary

A new arXiv paper (2605.15183) introduces tensor similarity, a weight-based metric for mechanistic interpretability that determines whether two networks or model components implement the same computation. Existing similarity measures rely either on empirical behavior, which is blind to out-of-distribution mechanisms, or on basis-dependent parameters, which are affected by weight-space symmetries. Tensor similarity addresses both issues for tensor-based models: it is invariant to weight-space symmetries, captures global functional equivalence, and accounts for cross-layer mechanisms via an efficient recursive algorithm. Empirically, the metric tracks functional training dynamics—including grokking and backdoor insertion—with higher fidelity than existing metrics. The work provides a prerequisite tool for verifying computational equivalence in interpretability research.

Paper Overview

Field: Machine Learning

Authors: ML Nissen Gonzalez, Melwina Albuquerque, Laurence Wroe, Jacob Meyer Cohen, Logan Riggs Smith, Thomas Dooms

Published: 2026-05-14

arXiv: 2605.15183

Abstract

Mechanistic interpretability aims to break models into meaningful parts; verifying that two such parts implement the same computation is a prerequisite. Existing similarity measures evaluate either empirical behaviour, leaving them blind to out-of-distribution mechanisms, or basis-dependent parameters, meaning they disregard weight-space symmetries.

To address these issues for the class of tensor-based models, the authors introduce a weight-based metric, tensor similarity, that is invariant to such symmetries. This metric captures global functional equivalence and accounts for cross-layer mechanisms using an efficient recursive algorithm.

Empirically, tensor similarity tracks functional training dynamics, such as grokking and backdoor insertion, with higher fidelity than existing metrics.

Key Points

  • Verifying computational equivalence between network components is a prerequisite for mechanistic interpretability
  • Behavior-based metrics miss out-of-distribution mechanisms; parameter-based metrics are sensitive to weight-space symmetries
  • Tensor similarity is a weight-based metric invariant to these symmetries
  • It captures global functional equivalence including cross-layer mechanisms via an efficient recursive algorithm
  • Empirically tracks grokking and backdoor insertion dynamics more faithfully than prior metrics
---

*Source: arXiv:2605.15183*

Tags

#mechanistic-interpretability#tensor-similarity#arxiv#machine-learning#model-equivalence#grokking#neural-networks

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620062