VECA: Elastic Attention Cores Bring Vision Transformers from O(N²) to O(N)
> *"If you can't simplify two things, you haven't found the simple way."* — paraphrasing Richard Feynman
This forum post is an in-depth Chinese-language explainer of the paper "Elastic Attention Cores for Scalable Vision Transformers" (Song, Chen & Nan, 2025), which introduces VECA — Visual Elastic Core Attention. Below is a structured English summary of its key arguments.
Key points
1. The problem: quadratic attention at high resolution
- Standard ViT treats an image as N patches and applies all-to-all self-attention, costing O(N²).
- Doubling resolution quadruples N but multiplies compute 16×. The post's table:
- This is why high-resolution domains (medical imaging, remote sensing, autonomous driving, industrial inspection) strain ViTs.
- VECA introduces C learnable core embeddings (C ≪ N). Patches attend only to cores; cores attend to each other. Information flows: Patch A → Core X → Core Y → Patch B.
- Complexity drops to O(N×C) + O(C²) ≈ O(N) when C is fixed. With C=64, theoretical savings reach ~256× at 2048×2048.
- Cores are randomly initialized and learned end-to-end — roles like aggregating sky, foreground objects, or contours emerge from training, not hand design.
- Cores persist across all Transformer layers, carrying and refining information layer by layer.
- Crucially, VECA keeps the full N patch tokens — unlike Perceiver or Set Transformer, which compress input through a small bottleneck. The paper: *"Compared to prior cross-attention architectures, VECA maintains and iteratively updates the full N input tokens, avoiding a small C-way bottleneck."* This preserves spatial detail needed for dense prediction.
- Classification (ImageNet): *"VECA achieves performance competitive with the latest vision foundation models while reducing computational cost."*
- Dense tasks (segmentation, detection): Remaining competitive confirms that keeping the full token sequence avoids the spatial-detail loss that cripples compressed-input methods.
- VECA is trained with nested training along the core axis, so at inference you can choose how many cores to use:
- Fewer cores → faster, cheaper (e.g., mobile/edge)
- More cores → higher accuracy (e.g., server-side)
- One set of weights, an adjustable "compute gearbox" — no retraining required for different budgets.
- VECA challenges a deep assumption in the Transformer community: that all-to-all pairwise interaction is necessary. The paper demonstrates that *"effective visual representations can be learned without any direct patch-to-patch interaction."*
- Its core-periphery topology mirrors efficient structures in nature and society (brain hubs, transit systems, social-network super-nodes) — a complex-systems insight imported into architecture design.
- Beyond raw speed, it opens a design space of alternative interaction topologies (graph, hierarchical, dynamic-routing attention).
- Song, A. Z., Chen, Y., & Nan, M. (2025). *Elastic Attention Cores for Scalable Vision Transformers.* arXiv preprint.
- Related: ViT (Dosovitskiy et al., 2020); Attention Is All You Need (Vaswani et al., 2017); Perceiver (Jaegle et al., 2021); Set Transformer (Lee et al., 2019); Linear Attention (Katharopoulos et al., 2020).
| Resolution | N (16×16 patches) | Attention cost | Relative | |-----------|-------------------|----------------|----------| | 224×224 | 196 | 38,416 | 1× | | 512×512 | 1,024 | 1,048,576 | 27× | | 1024×1024 | 4,096 | 16,777,216 | 437× | | 2048×2048 | 16,384 | 268,435,456 | 6,987× |