[论文] Backend-Agnostic Sparse Attention for Fast High-Resolution Visual Gene...

研究领域: CV 作者: Liao Ma, Jiayi Song, Yunfeng Wu, Songhua Liu, Peilin Zhao 发布时间: 2026-10-06 arXiv: 2610.08772

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Liao Ma, Jiayi Song, Yunfeng Wu, Songhua Liu, Peilin Zhao 发布时间: 2026-10-06 arXiv: 2610.08772

中文摘要

扩散Transformer(DiT)在图像和视频生成中取得了强大的性能,但全注意力的二次复杂度使高分辨率生成的计算成本高昂。窗口注意力提供了高效的替代方案,但现有方法面临实际的权衡:孤立的窗口阻止了跨窗口交互,通常在生成结果中引入可见的网格伪影;细粒度滑动窗口注意力虽然有效恢复相邻窗口之间的交互并提高视觉质量,但其不规则的计算模式在理论和实际加速之间造成了巨大差距,并需要针对每个硬件后端定制的专用内核。我们提出BASA,一种后端无关的稀疏注意力,兼得视觉质量和实际加速。具体而言,BASA用移位的局部窗口注意力替换视觉自注意力,通过在DiT块之间引入结构化的窗口移位方案,允许被窗口边界分隔的token在后续层中通信,从而实现全局信息交换并消除窗口引起的视觉伪影。我们的设计没有引入额外的不规则算子或定制内核,使其可直接部署在现有的注意力后端上。实验表明,BASA在FLUX上实现了超过理论估计90%的实测加速,在Wan上实现了4.52倍注意力加速,同时保持有竞争力的生成质量。

原文摘要

Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive. Window attention offers an efficient alternative, yet existing methods face a practical trade-off: partitioned window attention typically achieves computational efficiency consistent with its theoretical complexity. However, isolated windows block cross-window interaction, often introducing visible grid-like artifacts in the generated results. Fine-grained sliding-window attention effectively restores interactions across neighboring windows and improves visual quality. However, its irregular computation patterns create a substantial gap between theoretical and practical speedups and require s...


*自动采集于 2026-10-08*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens