论文概要
研究领域: CV
作者: Owais Iqbal, Sudipta Sarkar, Shyam Marjit
发布时间: 2026-09-30
arXiv: 2609.26768
中文摘要
我们提出 VideoMSN,一个用于高效自监督时空视频表示学习的掩码孪生网络框架。我们不再依赖沉重的 3D 架构或基于重建的自编码器来利用无标签数据学习,而是重新利用标准图像 Vision Transformer,将视频表示为超级图像——即从视频中采样帧组成的网格。从每个超级图像中,我们构建两个视图:一个进行空间块掩码,另一个进行时间帧掩码,确保帧间无信息泄漏。共享的 ViT 编码器使用掩码孪生损失对齐它们的嵌入,无需重建即可捕获运动和外观线索。我们无解码器的方案利用图像基础模型实现高效的视频表示学习。从预训练的 DINO-v3 和 DeiT-v3 图像编码器出发,VideoMSN 在 Kinetics-400、UCF101 和 HMDB51 上达到了最先进的性能,同时相比先前的视频自监督学习方法,视频预训练 epoch 减少了多达 32 倍和 160 倍。我们提出的方法在低样本分类中也表现出色,证实了所学表示在标签稀缺场景下的可迁移性。
原文摘要
We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards ...
自动采集于 2026-10-02
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。