[论文] One Block, Multiple Depths: Recurrent Vision Transformers with Depth-P...

研究领域: CV 作者: Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos 发布时间: 2026-10-08 arXiv: 2610.12448

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos 发布时间: 2026-10-08 arXiv: 2610.12448

中文摘要

本工作表明,单个 Transformer 块以循环方式应用即可在相当推理 FLOPs 下匹配全深度视觉编码器的精度,且无需中间特征蒸馏。reViT 通过将每个循环深度的 FFN 表示为一小组共享专家库的凸组合来恢复深度特异的变换。连续归一化深度坐标编程这个混合,定义一条可重采样的 FFN 参数空间轨迹。我们在两种设定下评估:ImageNet-1k 监督训练和自 DINOv2 教师蒸馏。在两种设定下,受控对比实验表明权重空间合并是测试过的最强 MoE 家族(在单 FFN 预算下),优于 token 调度和输出混合方案。从头训练时,reViT-B/16 以约 70% 更少的存储参数达到 DeiT III 的精度。仅用教师输出特征蒸馏的 8 专家模型几乎保留了 DINOv2 教师的全部线性探测精度,并迁移至分类、分割和深度预测任务。弹性深度训练允许一个检查点在多种测试深度下运行(重采样同一归一化坐标区间)。固定深度部署时,循环块可实例化为常规稠密图,消除在线路由与合并,计算量不变但部署存储增加。

原文摘要

In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained ...


*自动采集于 2026-10-10*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens