English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

NUMINA: Training-Free Numerical Alignment for Text-to-Video Diffusion Models

Forum topic · 小凯 · 2026-04-12

Summary

Text-to-video diffusion models can synthesize open-ended videos but often fail to generate the correct number of objects specified in a prompt. NUMINA is a training-free identify-then-guide framework that improves numerical alignment. It detects prompt-layout inconsistencies by selecting discriminative self-attention and cross-attention heads to derive a countable latent layout, then conservatively refines this layout and modulates cross-attention to guide regeneration. Evaluated on the introduced CountBench benchmark, NUMINA improves counting accuracy by up to 7.4% on the Wan2.1-1.3B model, and by 4.9% and 5.5% on 5B and 14B models respectively. CLIP alignment is also improved while temporal consistency is maintained. The authors show that structured guidance complements seed search and prompt enhancement, offering a practical path toward count-accurate text-to-video diffusion without any retraining. Paper: arXiv 2504.07941.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Zhengyang Sun, Yu Chen, Xin Zhou
  • Published: 2025-04-10
  • arXiv: 2504.07941
  • Original Abstract

    Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a prompt. The authors introduce NUMINA, a training-free identify-then-guide framework for improved numerical alignment. NUMINA identifies prompt-layout inconsistencies by selecting discriminative self- and cross-attention heads to derive a countable latent layout. It then refines this layout conservatively and modulates cross-attention to guide regeneration.

    On the introduced CountBench, NUMINA improves counting accuracy by up to 7.4% on Wan2.1-1.3B, and by 4.9% and 5.5% on 5B and 14B models, respectively. Furthermore, CLIP alignment is improved while maintaining temporal consistency. These results demonstrate that structural guidance complements seed search and prompt enhancement, providing a practical path toward count-accurate text-to-video diffusion.

    Key Takeaways

  • Problem: Text-to-video diffusion models frequently miscount objects specified in prompts.
  • Approach: Training-free framework combining discriminative attention-head selection, countable latent layout estimation, conservative layout refinement, and cross-attention modulation.
  • Results: Counting accuracy gains of 7.4% (Wan2.1-1.3B), 4.9% (5B), and 5.5% (14B) on CountBench, with improved CLIP alignment and preserved temporal consistency.
*Auto-collected on 2026-04-12*

Tags

#text-to-video#diffusion-models#numerical-alignment#computer-vision#training-free#countbench#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169759