English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising

Forum topic · 小凯 · 2026-06-21

Summary

JanusMesh (arXiv 2506.15890) is a training-free, text-driven framework for generating 3D visual illusions—single 3D meshes that reveal entirely different semantics from different viewpoints. Existing optimization-based methods are slow and produce oversaturated colors, while naive stitching approaches create geometrically inconsistent objects with visible seams and semantic leakage. JanusMesh addresses this with a two-stage pipeline: (1) a cross-space dual-branch denoising process that dynamically decodes 3D latents into voxel space for CLIP-guided directional alignment and signed distance field (SDF) blending, ensuring seamless geometric fusion; and (2) a view-conditioned texture synthesis module that projects and aggregates view-specific 2D diffusion priors onto the fused geometry. Experiments show the method generates highly realistic dual-semantic 3D illusions in only 3-5 minutes, significantly outperforming prior methods in geometric integrity, semantic recognizability, and efficiency.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Siang-Ling Zhang, Huai-Hsun Cheng, Tsung-Ju Yang
  • Published: 2026-06-19
  • arXiv: 2506.15890
  • Abstract

    Creating 3D visual illusions—a single 3D mesh that presents entirely different semantics when viewed from different angles—is a fascinating but highly challenging task.

    Existing approaches have notable drawbacks:

  • Optimization-based methods are slow and tend to produce overly saturated colors.
  • Simple stitching methods fail to generate geometrically consistent objects, resulting in visible unnatural seams and semantic leakage.
  • Method

    This paper proposes a fast, training-free, text-driven framework for 3D visual illusion generation, consisting of two stages:

    1. Cross-space dual-branch denoising: The 3D latents are dynamically decoded into voxel space, where CLIP-guided directional alignment and signed distance field (SDF) blending are performed to ensure seamless geometric fusion. 2. View-conditioned texture synthesis: A module that projects view-specific 2D diffusion priors and aggregates them onto the fused geometry.

    Results

    Extensive experiments demonstrate that the method generates highly realistic dual-semantic 3D illusions in just 3–5 minutes, significantly outperforming existing approaches in:

  • Geometric integrity
  • Semantic recognizability
  • Efficiency
---

*Automatically collected on 2026-06-21.*

Tags

#3d-visual-illusion#computer-vision#text-to-3d#diffusion-models#clip-guidance#sdf#zero-shot#janusmesh

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981600