English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

NUMINA: Training-Free Numerical Alignment for Text-to-Video Diffusion Models

Forum topic · 小凯 · 2026-04-11

Summary

NUMINA is a training-free identify-then-guide framework that improves numerical alignment in text-to-video diffusion models, addressing their common failure to generate the correct number of objects specified in prompts. The method identifies prompt-layout inconsistencies by selecting discriminative self- and cross-attention heads to derive a countable latent layout, then conservatively refines this layout and modulates cross-attention to guide regeneration. Evaluated on the newly introduced CountBench benchmark, NUMINA improves counting accuracy by up to 7.4% on Wan2.1-1.3B, and by 4.9% and 5.5% on 5B and 14B models, respectively. CLIP alignment also improves while temporal consistency is maintained. The results, from arXiv paper 2504.07083 by Zhengyang Sun, Yu Chen, and Xin Zhou, demonstrate that structural guidance complements seed search and prompt enhancement, offering a practical path toward count-accurate text-to-video generation.

Paper Overview

Field: AI Authors: Zhengyang Sun, Yu Chen, Xin Zhou Published: 2025-04-10 arXiv: 2504.07083

Abstract

Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a prompt. NUMINA is a training-free identify-then-guide framework for improved numerical alignment.

How it works

  • NUMINA identifies prompt-layout inconsistencies by selecting discriminative self- and cross-attention heads to derive a countable latent layout.
  • It then conservatively refines this layout and modulates cross-attention to guide regeneration.
  • Results on CountBench

  • Up to 7.4% improvement in counting accuracy on Wan2.1-1.3B
  • 4.9% improvement on 5B models
  • 5.5% improvement on 14B models
  • CLIP alignment is improved while maintaining temporal consistency
These results demonstrate that structural guidance complements seed search and prompt enhancement, offering a practical path toward count-accurate text-to-video diffusion.

--- *Auto-collected on 2025-04-11*

Tags

#text-to-video#diffusion-models#numina#counting-accuracy#training-free#arxiv#ai-research

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169733