论文概要
研究领域: CV
作者: Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos
发布时间: 2026-10-08
arXiv: 2610.12448
中文摘要
本工作表明,单个 Transformer 块以循环方式应用即可在相当推理 FLOPs 下匹配全深度视觉编码器的精度,且无需中间特征蒸馏。reViT 通过将每个循环深度的 FFN 表示为一小组共享专家库的凸组合来恢复深度特异的变换。连续归一化深度坐标编程这个混合,定义一条可重采样的 FFN 参数空间轨迹。我们在两种设定下评估:ImageNet-1k 监督训练和自 DINOv2 教师蒸馏。在两种设定下,受控对比实验表明权重空间合并是测试过的最强 MoE 家族(在单 FFN 预算下),优于 token 调度和输出混合方案。从头训练时,reViT-B/16 以约 70% 更少的存储参数达到 DeiT III 的精度。仅用教师输出特征蒸馏的 8 专家模型几乎保留了 DINOv2 教师的全部线性探测精度,并迁移至分类、分割和深度预测任务。弹性深度训练允许一个检查点在多种测试深度下运行(重采样同一归一化坐标区间)。固定深度部署时,循环块可实例化为常规稠密图,消除在线路由与合并,计算量不变但部署存储增加。
原文摘要
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained ...
自动采集于 2026-10-10
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。