Paper Overview
Research Area: Computer Vision (CV) Authors: Yen-Siang Wu, Rundong Luo, Jingsen Zhu, Tao Tu, Ali Farhadi, Matthew Wallingford, Yu-Chiang Frank Wang, Steve Marschner, Wei-Chiu Ma Published: 2026-04-23 arXiv: 2604.21931
Abstract
How can we tell whether a video has been sped up or slowed down? How can we generate videos at different speeds? Although videos have been central to modern computer vision research, little attention has been paid to perceiving and controlling the passage of time.
In this paper, the authors study time as a learnable visual concept and develop models for reasoning about and manipulating the flow of time in videos.
Key Contributions
- Temporal reasoning via self-supervision: The models exploit multimodal cues and temporal structure naturally present in videos to learn, in a self-supervised manner, to detect speed changes and estimate playback speed.
- Largest slow-motion dataset to date: The learned temporal reasoning models enable curating the largest slow-motion video dataset to date from noisy in-the-wild sources. Such slow-motion footage, typically filmed by high-speed cameras, contains richer temporal detail than ordinary video.
- Speed-conditioned video generation: Models that generate motion at specified playback speeds, enabling direct control over the temporal flow of generated video.
- Temporal super-resolution: Models that convert low-frame-rate, blurry videos into high-frame-rate sequences with fine-grained temporal detail.
Takeaways
The findings highlight that time is a manipulable perceptual dimension in video learning, opening doors to temporally controllable video generation, temporal forensics detection, and potentially richer world models that understand how events unfold over time.
--- *Originally posted on zhichai.net, auto-collected 2026-04-25.*