论文概要
研究领域: CV
作者: Neel Varma, Andrew Rufail, Dipika Khullar, Vasu Sharma
发布时间: 2026-10-02
arXiv: 2610.03698
中文摘要
自监督视觉 Transformer(如 DINOv2)能学到丰富的视觉表示,但其内部 token 的功能仍不为人所理解。最近的架构引入专门的寄存器 token 来减少在背景区域出现的高范数离群 patch token,但这两类 token 的语义和功能角色尚未完全确立。本文通过在 DINOv2 的寄存器 token 和离群 token 激活上训练稀疏自编码器(SAE)来分析其角色。通过自动化可解释性流程、UMAP 聚类和 CLIP 空间交叉验证,我们发现寄存器 token 特征与高层语义概念的关联更强,而离群 token 特征更多与低层结构、背景和纹理主导模式相关。因果消融实验进一步揭示了显著的功能不对称性:破坏激活最强的寄存器衍生特征会导致表示余弦相似度下降 48.17%,而破坏离群衍生特征仅导致 0.31% 的下降。这些结果为自监督 ViT 中的 token 特化提供了证据。
原文摘要
Self-supervised Vision Transformers (ViTs), such as DINOv2, learn rich visual representations, but the functions of their internal tokens remain poorly understood. Recent architectures introduce dedicated register tokens to reduce high-norm out- lier patch tokens that emerge in background re- gions, yet the semantic and functional roles of both token types have not been fully established. In this paper, we analyze these roles by training sparse autoencoders (SAEs) on register-token and outlier-token activations in DINOv2. Using an automated interpretability pipeline, UMAP clus- tering, and CLIP-space cross-checks, we find that register-token features are more strongly associ- ated with high-level semantic concepts. Outlier- token features, by contrast, are more often associ- ated with lowe...
自动采集于 2026-10-06
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。