English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AnimationBench: A Benchmark for Evaluating Character-Centric Animation in Image-to-Video Generation

Forum topic · 小凯 · 2026-04-18

Summary

AnimationBench (arXiv:2504.13082) is the first systematic benchmark for evaluating animation-style image-to-video (I2V) generation. Existing benchmarks, designed mainly for realistic videos, struggle to assess stylized appearance, exaggerated motion, and character-centric consistency. AnimationBench operationalizes the Twelve Basic Principles of Animation and IP Preservation into measurable evaluation dimensions, complemented by broader quality dimensions such as semantic consistency, motion plausibility, and camera-motion consistency. It supports both standardized closed-set evaluation for reproducible comparisons and flexible open-set evaluation for diagnostic analysis, using vision-language models for scalable scoring. Experiments show the benchmark aligns well with human judgment and reveals animation-specific quality gaps that realism-oriented benchmarks miss.

Overview

  • Field: Computer Vision
  • Authors: Leyi Wu, Pengjun Fang, Kai Sun
  • Published: 2025-04-17
  • arXiv: 2504.13082
  • Abstract

    Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks—largely designed for realistic videos—struggle to evaluate animation-style generation with its stylized appearance, exaggerated motion, and character-centric consistency. Moreover, they rely on fixed prompt sets and rigid pipelines, offering limited flexibility for open-domain content and custom evaluation needs.

    To address this gap, the authors introduce AnimationBench, the first systematic benchmark for evaluating animation image-to-video generation. Key contributions:

  • Operationalizes the Twelve Basic Principles of Animation and IP Preservation into measurable evaluation dimensions.
  • Includes Broader Quality Dimensions such as semantic consistency, motion plausibility, and camera-motion consistency.
  • Supports both standardized closed-set evaluation (for reproducible comparisons) and flexible open-set evaluation (for diagnostic analysis).
  • Leverages vision-language models for scalable, automated assessment.
Extensive experiments show that AnimationBench aligns closely with human judgments and reveals animation-specific quality differences overlooked by realism-oriented benchmarks, enabling more informative and discriminative evaluation of state-of-the-art I2V models.

---

*Auto-collected on 2026-04-18*

Tags

#animation#video-generation#benchmark#image-to-video#computer-vision#evaluation#vision-language-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618543