English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Seeing Fast and Slow: Learning the Flow of Time in Videos

Forum topic · 小凯 · 2026-04-25

Summary

This paper studies time as a learnable visual concept in computer vision, developing models that can reason about and manipulate the flow of time in videos. The authors first learn, in a self-supervised manner, to detect speed changes and estimate playback speed by exploiting multimodal cues and temporal structure naturally present in videos. These learned temporal reasoning models are then used to curate the largest slow-motion video dataset to date from noisy in-the-wild sources. Leveraging this data, the team develops models enabling temporal control, including speed-conditioned video generation that produces motion at specified playback speeds, and temporal super-resolution that converts low-frame-rate, blurry videos into high-frame-rate sequences with fine-grained temporal detail. The work demonstrates that time is a manipulable perceptual dimension in video learning, opening doors to temporally controllable video generation, temporal forensics, and richer world models that understand how events unfold over time. Authors include researchers from Cornell and the University of Washington. arXiv: 2604.21931.

Paper Overview

Research Area: Computer Vision (CV) Authors: Yen-Siang Wu, Rundong Luo, Jingsen Zhu, Tao Tu, Ali Farhadi, Matthew Wallingford, Yu-Chiang Frank Wang, Steve Marschner, Wei-Chiu Ma Published: 2026-04-23 arXiv: 2604.21931

Abstract

How can we tell whether a video has been sped up or slowed down? How can we generate videos at different speeds? Although videos have been central to modern computer vision research, little attention has been paid to perceiving and controlling the passage of time.

In this paper, the authors study time as a learnable visual concept and develop models for reasoning about and manipulating the flow of time in videos.

Key Contributions

  • Temporal reasoning via self-supervision: The models exploit multimodal cues and temporal structure naturally present in videos to learn, in a self-supervised manner, to detect speed changes and estimate playback speed.
  • Largest slow-motion dataset to date: The learned temporal reasoning models enable curating the largest slow-motion video dataset to date from noisy in-the-wild sources. Such slow-motion footage, typically filmed by high-speed cameras, contains richer temporal detail than ordinary video.
  • Speed-conditioned video generation: Models that generate motion at specified playback speeds, enabling direct control over the temporal flow of generated video.
  • Temporal super-resolution: Models that convert low-frame-rate, blurry videos into high-frame-rate sequences with fine-grained temporal detail.

Takeaways

The findings highlight that time is a manipulable perceptual dimension in video learning, opening doors to temporally controllable video generation, temporal forensics detection, and potentially richer world models that understand how events unfold over time.

--- *Originally posted on zhichai.net, auto-collected 2026-04-25.*

Tags

#computer-vision#video-generation#slow-motion#temporal-super-resolution#self-supervised-learning#arxiv#world-models#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618729