English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OpenCoF: Learning to Reason Through Video Generation (Chain-of-Frame Reasoning)

Forum topic · 小凯 · 2026-07-12

Summary

OpenCoF is a research framework exploring Chain-of-Frame (CoF) reasoning, where video generation models reason through temporally connected frames instead of textual Chain-of-Thought (CoT). The authors from CUHK introduce two components: OpenCoF-17K, a reasoning video dataset spanning 11 task families, and Wan-CoF, a fine-tuned video generation model built to test whether diverse temporal supervision improves CoF behavior. On four video reasoning benchmarks, Wan-CoF delivers significant gains over the Wan2.2-I2V-A14B baseline. The work further explores advanced CoF capabilities by equipping the model with visual and textual reasoning tokens that capture low-level visual cues and high-level semantic priors for spatial and temporal reasoning, validated through performance comparisons and attention analysis across model depth, denoising steps, and spatial-temporal dimensions. Findings indicate strong video reasoning requires broad temporal supervision plus explicit mechanisms for organizing intermediate reasoning states. Dataset, model, and code are open-sourced. Paper: arXiv 2607.08763.

Paper Overview

Field: Computer Vision (CV) Authors: Xinyan Chen, Ziyu Guo, Renrui Zhang, Dongzhi Jiang, Hongsheng Li arXiv: 2607.08763

Summary

Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct from previous Chain-of-Thought (CoT): reasoning can unfold through temporally connected frames, known as Chain-of-Frame (CoF) reasoning. However, existing video generators are primarily trained on general video corpora, still lacking diverse supervision and dedicated designs for CoF reasoning.

To address this gap, the authors introduce OpenCoF, a framework comprising:

  • OpenCoF-17K: a reasoning video dataset spanning 11 task families
  • Wan-CoF: a fine-tuned video model for studying whether diverse temporal supervision improves CoF behavior
  • Across four video reasoning benchmarks, Wan-CoF achieves significant gains over the Wan2.2-I2V-A14B baseline.

    Advanced CoF Design

    Building on these results, the paper empirically explores more advanced CoF capabilities: equipping the model with visual and textual reasoning tokens. These tokens capture:

  • Low-level visual cues (visual tokens)
  • High-level semantic priors (textual tokens)
Both serve spatial and temporal reasoning. Through performance comparisons and attention analysis, the authors examine how these tokens contribute across model depth, denoising steps, and spatial-temporal dimensions.

Key Findings

Stronger video reasoning requires:

1. Broad, diverse temporal supervision 2. Explicit mechanisms for organizing intermediate reasoning states

The dataset, model, and code have been open-sourced to facilitate reasoning-oriented video generation research.

--- *Auto-collected on 2026-07-12*

Tags

#opencof#video-generation#chain-of-frame#reasoning#arxiv#computer-vision#fine-tuning#dataset

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178379392