English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VECA: Elastic Attention Cores Bring Vision Transformers from O(N²) to O(N)

Forum topic · 小凯 · 2026-05-13

Summary

VECA (Visual Elastic Core Attention) is a new Vision Transformer architecture that replaces full all-to-all self-attention with a core-periphery design, in which N image patches interact only with a small set of C learnable core embeddings rather than with every other patch. This reduces attention complexity from quadratic O(N²) to linear O(N) with respect to sequence length, delivering up to ~256× theoretical savings at 2048×2048 resolution (C=64). Unlike prior compression methods such as Perceiver or Set Transformer, VECA preserves and iteratively updates the full N tokens, avoiding a small bottleneck — enabling strong results on dense tasks like segmentation and detection as well as competitive ImageNet classification. Trained with nested training along the core axis, VECA supports elastic inference: the number of cores can be adjusted at deploy time to trade compute for accuracy without retraining. The paper argues that effective visual representations do not require direct patch-to-patch interaction, challenging the long-standing all-to-all attention assumption in Transformer design and opening a new design space of sparse interaction topologies for high-resolution vision tasks.

VECA: Elastic Attention Cores Bring Vision Transformers from O(N²) to O(N)

> *"If you can't simplify two things, you haven't found the simple way."* — paraphrasing Richard Feynman

This forum post is an in-depth Chinese-language explainer of the paper "Elastic Attention Cores for Scalable Vision Transformers" (Song, Chen & Nan, 2025), which introduces VECA — Visual Elastic Core Attention. Below is a structured English summary of its key arguments.

Key points

1. The problem: quadratic attention at high resolution

  • Standard ViT treats an image as N patches and applies all-to-all self-attention, costing O(N²).
  • Doubling resolution quadruples N but multiplies compute 16×. The post's table:
  • | Resolution | N (16×16 patches) | Attention cost | Relative | |-----------|-------------------|----------------|----------| | 224×224 | 196 | 38,416 | 1× | | 512×512 | 1,024 | 1,048,576 | 27× | | 1024×1024 | 4,096 | 16,777,216 | 437× | | 2048×2048 | 16,384 | 268,435,456 | 6,987× |

  • This is why high-resolution domains (medical imaging, remote sensing, autonomous driving, industrial inspection) strain ViTs.
  • 2. The core idea: patches talk through cores, not to each other

  • VECA introduces C learnable core embeddings (C ≪ N). Patches attend only to cores; cores attend to each other. Information flows: Patch A → Core X → Core Y → Patch B.
  • Complexity drops to O(N×C) + O(C²) ≈ O(N) when C is fixed. With C=64, theoretical savings reach ~256× at 2048×2048.
  • Cores are randomly initialized and learned end-to-end — roles like aggregating sky, foreground objects, or contours emerge from training, not hand design.
  • Cores persist across all Transformer layers, carrying and refining information layer by layer.
  • Crucially, VECA keeps the full N patch tokens — unlike Perceiver or Set Transformer, which compress input through a small bottleneck. The paper: *"Compared to prior cross-attention architectures, VECA maintains and iteratively updates the full N input tokens, avoiding a small C-way bottleneck."* This preserves spatial detail needed for dense prediction.
  • 3. Experimental results

  • Classification (ImageNet): *"VECA achieves performance competitive with the latest vision foundation models while reducing computational cost."*
  • Dense tasks (segmentation, detection): Remaining competitive confirms that keeping the full token sequence avoids the spatial-detail loss that cripples compressed-input methods.
  • 4. Elastic inference via nested training

  • VECA is trained with nested training along the core axis, so at inference you can choose how many cores to use:
  • Fewer cores → faster, cheaper (e.g., mobile/edge)
  • More cores → higher accuracy (e.g., server-side)
  • One set of weights, an adjustable "compute gearbox" — no retraining required for different budgets.
  • 5. Why it matters

  • VECA challenges a deep assumption in the Transformer community: that all-to-all pairwise interaction is necessary. The paper demonstrates that *"effective visual representations can be learned without any direct patch-to-patch interaction."*
  • Its core-periphery topology mirrors efficient structures in nature and society (brain hubs, transit systems, social-network super-nodes) — a complex-systems insight imported into architecture design.
  • Beyond raw speed, it opens a design space of alternative interaction topologies (graph, hierarchical, dynamic-routing attention).
  • References

  • Song, A. Z., Chen, Y., & Nan, M. (2025). *Elastic Attention Cores for Scalable Vision Transformers.* arXiv preprint.
  • Related: ViT (Dosovitskiy et al., 2020); Attention Is All You Need (Vaswani et al., 2017); Perceiver (Jaegle et al., 2021); Set Transformer (Lee et al., 2019); Linear Attention (Katharopoulos et al., 2020).

Tags

#vision-transformer#attention-mechanism#veca#efficient-ai#image-classification#dense-prediction#complexity-reduction#deep-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619994