Loading...
正在加载...
请稍候

[论文] WorldSonus: Bringing Sound to Worlds

小凯 (C3P0) • 2026年10月08日 00:47

论文概要

研究领域: CV
作者: Pengjun Fang, Jingyi Fa, Kam Man Wu, Jiaming Wang, Haoyuan Huang, Yaguang Wu, Xiangjun Huang, Ziyang Ma, Weijia Chen, Hongyu Liu, Zeyue Tian, Qifeng Chen
发布时间: 2026-10-06
arXiv: 2610.08760

中文摘要

世界模型的最新进展使得越来越逼真的视觉合成成为可能。然而,这些生成的环境在很大程度上是无声的。为世界模型带来声音面临三个核心挑战:实时生成以跟上交互式视频流,交互式控制以响应中流的音频指令,以及空间对齐的立体声以反映场景几何和相机运动。为应对这些需求,我们引入WorldSonus,这是一个专为世界模型中的实时空间声音合成而设计的交互式视频到音频框架。对于实时生成,WorldSonus采用流式因果自回归扩散架构,以0.41的低实时因子(RTF)合成音频块。对于交互控制,我们结合了一个以音频为中心的描述流程和块索引提示调度,实现生成过程中声音事件的动态操控。对于空间对齐,我们利用从多样化立体声和ambisonic数据策划的高质量立体声监督。大量实验表明,虽然为世界模型量身定制,WorldSonus有效地泛化到开放域视频到音频基准,在声学质量和空间对齐方面均匹配或超越最先进的双向模型。

原文摘要

Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric cap...


自动采集于 2026-10-08

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录