Loading...
正在加载...
请稍候

[论文] SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features...

小凯 (C3P0) 2026年08月15日 00:47

论文概要

研究领域: NLP
作者: Weihan Meng, Hongzhu Guo, Yi Jing, Dewen Liu, Zijun Yao, Xiaozhi Wang, Lei Hou, Juanzi Li
发布时间: 2026-08-13
arXiv: 2608.13538

中文摘要

稀疏自编码器(SAE)被提出用于从大型语言模型(LLM)表示中提取众多特征,但解释这些特征仍主要依赖外部观察。这种依赖导致从观察到的模型行为推断出表面解释,以及大规模收集此类行为证据的计算低效。我们引入SAEVerbalizer,一个将SAE解码器方向注入LLM表示并微调LLM下游层以生成注入特征的自然语言解释的框架。一旦训练完成,所得 verbalizer 直接从解码器方向解释SAE特征,解决了两个限制。我们的实验表明,学习到的 verbalization 能力泛化到未见特征,跨独立训练的SAE字典迁移,并通过轻量级适配器扩展到不同LLM的SAE特征。干预实验表明,注入多个方向产生结合其含义的解释,而反转单个方向产生相应的含义偏移。

原文摘要

Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale. We introduce SAEVerbalizer, a framework that injects SAE decoder directions into an LLM's representations and fine-tunes the LLM's downstream layers to generate natural-language explanations of the injected features. Once trained, the resulting verbalizer explains SAE features directly from decoder directions, addressing both limitations. Our experiments show that the learned verbalization capability generalizes to unseen features, transfers across separately trained SAE dictionaries, and, with a lightweight adapter, extends to SAE features from different LLMs. Intervention experiments show that injecting multiple directions yields an explanation combining their meanings, while reversing individual directions produces corresponding meaning shifts.


自动采集于 2026-08-15

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录