论文概要
研究领域: CV
作者: Tianjiao Yu, Xinzhuo Li, Yifan Shen, Ying Shen, Kiet A. Nguyen, Adheesh Sunil Juvekar, Ismini Lourentzou
发布时间: 2026-10-01
arXiv: 2610.02201
中文摘要
高分辨率 3D 生成越来越依赖体素隐空间和多阶段流水线——先预测活跃结构再合成局部几何。虽然有效,但这种设计将连续表面碎片化为大量局部 token,推高生成成本,且常削弱细长或高连接形状的结构一致性。我们提出 SILSA,一个拓扑感知的 3D 生成框架,使用紧凑的滑窗切片隐空间表示形状。SILSA 不生成昂贵的体素 token,而是沿三条正交轴使用固定数量的重叠切片,每个 token 概括一个局部深度窗口,保持横截面连续性并支持单阶段整流流生成。Slice VAE 将定向表面样本编码为多轴切片隐空间并用稀疏体解码器重建,Volumetric Anchor Lattice 通过共享 3D 工作空间协调定向切片流。为保持结构正确性,我们引入切片级拓扑监督,匹配持续图并对齐相邻切片间的 Betti 数转移。实验表明 SILSA 显著提升结构保真度并大幅降低生成成本:PSNR 提升 8.7%,覆盖率提升 5.96 个百分点,Betti 误差降低 9.2%(均优于最强基线),同时比次紧凑基线少用 70% token,比稀疏或分层 tokenizer 少用超 98% token,有效降低训练内存 40.4%、推理时间 58.5%。定性结果进一步显示对薄结构、重复组件和长程连通性的保留均有改善。
原文摘要
High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-a...
自动采集于 2026-10-03
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。