Paper Overview
Field: Machine Learning Authors: Nikolai Bolik, Lennart Stöpler, Artur Andrzejak Published: 2026-08-11 arXiv: 2608.11197
Abstract (Full Translation)
Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure.
We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts analysis to SAE set similarities, we find that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states; rather, they track the model's internal similarity structure.
To further probe this gap, we study active latent sets under controlled semantic modifications, revealing a notable mismatch between human judgments of concept change and changes in SAE active sets. We interpret this as evidence that, beyond idealized settings, SAE features do not compose through simple bag-of-features semantics.
Key Takeaways
- Set-level similarity: The authors propose overlap over active SAE latent sets as an interpretable alternative to cosine similarity on dense embeddings.
- Validated methodology: The measure recovers compositional (union-like) structure in toy models and yields semantically coherent neighborhoods in natural text.
- Negative result: SAE activation sets are no more faithful to human category boundaries or typicality than dense embeddings or residual-stream states.
- Instability under modification: Active latent sets change in ways that diverge significantly from human judgments of semantic change.
- Conclusion: SAE features should not be assumed to compose as a simple "bag of features" outside idealized conditions.