Summary
AnimationBench (arXiv:2504.13082) is the first systematic benchmark for evaluating animation-style image-to-video (I2V) generation. Existing benchmarks, designed mainly for realistic videos, struggle to assess stylized appearance, exaggerated motion, and character-centric consistency. AnimationBench operationalizes the Twelve Basic Principles of Animation and IP Preservation into measurable evaluation dimensions, complemented by broader quality dimensions such as semantic consistency, motion plausibility, and camera-motion consistency. It supports both standardized closed-set evaluation for reproducible comparisons and flexible open-set evaluation for diagnostic analysis, using vision-language models for scalable scoring. Experiments show the benchmark aligns well with human judgment and reveals animation-specific quality gaps that realism-oriented benchmarks miss.
Overview
- Field: Computer Vision
- Authors: Leyi Wu, Pengjun Fang, Kai Sun
- Published: 2025-04-17
- arXiv: 2504.13082
Abstract
Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks—largely designed for realistic videos—struggle to evaluate animation-style generation with its stylized appearance, exaggerated motion, and character-centric consistency. Moreover, they rely on fixed prompt sets and rigid pipelines, offering limited flexibility for open-domain content and custom evaluation needs.
To address this gap, the authors introduce AnimationBench, the first systematic benchmark for evaluating animation image-to-video generation. Key contributions:
- Operationalizes the Twelve Basic Principles of Animation and IP Preservation into measurable evaluation dimensions.
- Includes Broader Quality Dimensions such as semantic consistency, motion plausibility, and camera-motion consistency.
- Supports both standardized closed-set evaluation (for reproducible comparisons) and flexible open-set evaluation (for diagnostic analysis).
- Leverages vision-language models for scalable, automated assessment.
Extensive experiments show that AnimationBench aligns closely with human judgments and reveals animation-specific quality differences overlooked by realism-oriented benchmarks, enabling more informative and discriminative evaluation of state-of-the-art I2V models.
---
*Auto-collected on 2026-04-18*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177618543