Summary
This paper examines whether the Compressed Computation (CC) toy model of Braun et al. (2025) is truly an instance of computation in superposition. The CC model appears to compute 100 ReLU functions using only 50 neurons, achieving better loss than expected from representing just 50 ReLU functions. The authors show that the model mixes inputs through its noisy residual stream, corresponding to an unintended mixing matrix in the labels. Decomposing the training objective into a ReLU term and a mixing term, they find that performance gains scale with the magnitude of the mixing matrix and disappear when the matrix is removed. Learned neuron directions concentrate in the subspace associated with the top 50 eigenvalues of the mixing matrix, indicating the mixing term governs the solution. Additionally, a semi-non-negative matrix factorization (SNMF) baseline derived solely from the mixing matrix reproduces the qualitative loss curve and improves on prior baselines, though it does not fully match the trained model. These results suggest that CC is not an appropriate toy model for computation in superposition. Paper: arXiv 2606.14673.
Paper Overview
Field: Machine Learning
Authors: Jai Bhagat, Sara Molas-Medina, Giorgi Giglemiani
Published: 2026-06-12
arXiv: 2606.14673
Abstract
We study whether the Compressed Computation (CC) toy model (Braun et al., 2025) is an instance of computation in superposition. The CC model appears to compute 100 ReLU functions with just 50 neurons, achieving a better loss than expected from only representing 50 ReLU functions. We show that the model mixes inputs via its noisy residual stream, corresponding to an unintended mixing matrix in the labels. Splitting the training objective into the ReLU term and the mixing term, we find that performance gains scale with the magnitude of the mixing matrix and vanish when the matrix is removed. The learned neuron directions concentrate in the subspace associated with the top 50 eigenvalues of the mixing matrix, suggesting that the mixing term governs the solution. Finally, a semi-non-negative matrix factorization (SNMF) baseline derived only from the mixing matrix reproduces the qualitative loss curve and improves on previous baselines, although it does not match the trained model. These results indicate that CC is not a suitable toy model for computation in superposition.
---
*Auto-collected on 2026-06-16*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177981392