[论文] Common-Mode Collapse and Recovery in Direct Feedback Alignment
研究领域: ML 作者: Varun Reddy, Bernardo L. Sabatini, Houman Safaai 发布时间: 2026-09-25 arXiv: 2609.31589
论文概要
研究领域: ML 作者: Varun Reddy, Bernardo L. Sabatini, Houman Safaai 发布时间: 2026-09-25 arXiv: 2609.31589
中文摘要
直接反馈对齐(DFA)通过输出误差的固定随机投影来训练隐藏层。在 tanh 隐藏单元加独立 sigmoid 输出的设定下,朴素的随机梯度下降可能停滞在“用类别频率常数预测”对应的损失水平附近。我们将这种停滞归因于误差的共模分量——即跨输入共享的分量。精确的均值—协方差分解分离出一个由平均教学信号与平均突触前活动构成的秩一更新,其主导分量把 tanh 单元推向饱和。在初始化时,随机反馈平均而言无法系统性地纠正共享误差;读出层的学习限制了停滞的持续时间。一个从网络初始化、不含拟合参数的简化模型,能在 48 种设定下预测激活敏感度的集中程度。在 MNIST 上,类别可解码性在崩塌后大体幸存,但固定学习率下读出层学习仍然缓慢;Adam 尽管崩塌更深却学得更快。将基线读出层校准到类别先验可以抑制崩塌并加速学习;更弱的反馈则以更慢的学习换取更少的崩塌。用误差符号替代误差会维持崩塌;减去信号的批次均值则可防止持续性崩塌并改善测试设定下的学习。类似效应也出现在更深的网络、卷积网络和 CIFAR-10 上,其严重程度与代价取决于读出层、优化器和输入统计特性。
原文摘要
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading component drives tanh units toward saturation. At initialization, random feedback provides no systematic correction of the shared error on average; readout learning limits its duration. A reduced model initialized from the network, without fitted parameters, predicts the concentration of activation sensitiv...
*自动采集于 2026-09-29*
#论文 #arXiv #ML #小凯