English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LoViF 2026 PhyScore Challenge: Holistic Quality Assessment for 4D World Model-Generated Videos

Forum topic · 小凯 · 2026-05-08

Summary

The LoViF 2026 PhyScore challenge addresses a key gap in evaluating world-model-generated videos: perceptual quality alone cannot determine whether generated dynamics are physically plausible, temporally coherent, and consistent with input conditions. The competition requires participants to build metrics that jointly predict four dimensions—video quality, physical realism, condition-video alignment, and temporal consistency—and to localize physical anomaly timestamps for fine-grained diagnosis. The benchmark contains 1,554 videos generated by seven representative world generative models, organized into three tracks (text-to-2D, image-to-4D, and video-to-4D) spanning 26 categories that cover physics-relevant scenarios such as dynamics, optics, and thermodynamics, plus diverse real-world and creative content. Labels were produced through trained human annotation with an automated quality-control pass. Evaluation combines score prediction and anomaly localization using a composite protocol based on TimeStamp_IOU and SRCC/PLCC. This paper summarizes the challenge design and offers method-level insights from submitted solutions.

Overview

This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and 4D generation settings.

Motivation

The challenge addresses a central gap in current evaluation practice: perceptual quality alone is insufficient to judge whether generated dynamics are physically plausible, temporally coherent, and consistent with input conditions.

Task Design

Participants are required to:

  • Build a metric that jointly predicts four dimensions:
  • Video Quality
  • Physical Realism
  • Condition-Video Alignment
  • Temporal Consistency
  • Localize physical anomaly timestamps for fine-grained diagnosis.
  • Benchmark Dataset

  • 1,554 videos generated by seven representative world generative models
  • Three tracks: text-to-2D, image-to-4D, and video-to-4D
  • 26 categories explicitly covering physics-relevant scenarios (dynamics, optics, thermodynamics) as well as diverse real-world and creative content
  • Labels (scores and anomaly timestamps) were produced through trained human annotation, supplemented with an additional automated quality-control pass to ensure reliability.

    Evaluation Protocol

    Evaluation is based on both score prediction and anomaly localization, using a composite protocol that combines TimeStamp_IOU with SRCC/PLCC.

    The report summarizes the challenge design and provides method-level insights from the submitted solutions.

    References

  • arXiv: 2605.05187

Tags

#world-models#video-quality-assessment#4d-generation#physics-realism#benchmark#computer-vision#challenge#temporal-consistency

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619586