English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling

Forum topic · 小凯 · 2026-07-01

Summary

LeVo 2 is a hybrid LLM-Diffusion framework for controllable full-length song generation, presented in arXiv paper 2507.00002 by Shun Lei, Huaicheng Zhang, and Dapeng Wu. Existing language model-based song generation systems face a structural trade-off: mixed-token modeling preserves vocal-instrument coordination but obscures track-specific details, while dual-track prediction improves acoustics but weakens global planning. LeVo 2 resolves this via hierarchical modeling: a language model (LeLM) first predicts mixed tokens for semantic planning, then predicts vocal and accompaniment tokens in parallel for track-specific refinement, with a diffusion-based Music Codec reconstructing full-length waveforms. The extended version introduces an aesthetics-guided alignment training scheme: automated music aesthetics evaluation assigns musicality-tier conditions during pretraining, followed by progressive post-training with SFT, large-scale offline DPO, and closed-loop semi-online DPO to improve quality, controllability, and musicality. Modular extension then trains track-specific LMs for acoustic refinement. Expert listening tests and objective evaluations show LeVo 2 outperforms open-source baselines on six subjective dimensions and approaches leading commercial systems on several listening metrics.

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling

  • Research area: Audio generation
  • Authors: Shun Lei, Huaicheng Zhang, Dapeng Wu
  • arXiv: 2507.00002
  • Overview

    Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics and prompts. Existing language model-based systems face a structural trade-off: mixed-token modeling preserves vocal-instrument coordination but obscures track-specific details, whereas dual-track prediction improves acoustics but requires longer sequences and weakens global planning.

    The authors present LeVo 2, a hybrid LLM-Diffusion framework for controllable full-length song generation. LeVo 2 formulates this trade-off as hierarchical modeling:

    1. LeLM first predicts mixed tokens for semantic planning, then predicts vocal and accompaniment tokens in parallel for track-specific refinement. 2. A diffusion-based Music Codec reconstructs full-length waveforms.

    Key Contributions (Extended Version)

    The core contribution of this extended version is an aesthetics-guided alignment training scheme:

  • Pretraining: An automated music aesthetics evaluation framework assigns musicality-tier conditions to large-scale data, providing musicality priors before preference alignment.
  • Progressive post-training: SFT, large-scale offline DPO, and closed-loop semi-online DPO successively improve generation quality, controllability, and musicality.
  • Modular extension: Track-specific LMs are trained for acoustic refinement while preserving the aligned semantic planner.
This scheme separates musicality learning, controllability alignment, and acoustic refinement, mitigating optimization conflicts and the limitations of static offline preference pairs.

Results

Expert listening tests and objective evaluations show that LeVo 2 outperforms open-source baselines across six subjective dimensions and approaches leading commercial systems on several listening metrics.

---

*Auto-collected on 2026-07-01.*

Tags

#paper#audio-generation#music-generation#llm#diffusion#song-generation#alignment#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208337