English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

JanusMesh: Fast, Training-Free 3D Visual Illusion Generation via Cross-Space Dual-Branch Denoising

Forum topic · 小凯 · 2026-06-20

Summary

JanusMesh (arXiv:2506.16809) is a fast, training-free framework for text-driven 3D visual illusion generation—creating a single 3D mesh that reveals entirely different semantics from different viewing angles. Existing optimization-based methods are slow and prone to oversaturated colors, while naive stitching produces geometrically incoherent objects with visible seams and semantic leaks. JanusMesh decouples generation into two stages: (1) a cross-space dual-branch denoising process that dynamically decodes 3D latents into voxel space for CLIP-guided orientation alignment and Signed Distance Field (SDF) blending, ensuring seamless geometric fusion; and (2) a viewpoint-conditioned texture synthesis module that projects and aggregates viewpoint-specific 2D diffusion priors onto the blended geometry. Experiments show the method generates highly realistic dual-semantic 3D illusions in only 3-5 minutes, significantly outperforming prior methods in geometric integrity, semantic recognizability, and efficiency. Authored by Siang-Ling Zhang, Huai-Hsun Cheng, and Tsung-Ju Yang.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Siang-Ling Zhang, Huai-Hsun Cheng, Tsung-Ju Yang
  • Published: 2025-06-20
  • arXiv: 2506.16809
  • Abstract

    Creating 3D visual illusions—a single 3D mesh that reveals entirely different semantics from various viewing angles—is a fascinating but tough challenge. Existing optimization-based methods are slow and can produce oversaturated colors. In contrast, naive stitching approaches fail to produce geometrically coherent objects, resulting in visible unnatural seams and semantic leaks.

    This paper presents a fast and training-free framework for generating text-driven 3D visual illusions. The approach decouples the generation into two stages:

    1. Cross-space dual-branch denoising process: dynamically decodes 3D latents into voxel space for CLIP-guided orientation alignment and Signed Distance Field (SDF) blending, which ensures seamless geometric fusion. 2. Viewpoint-conditioned texture synthesis module: projects and aggregates viewpoint-specific 2D diffusion priors onto the blended geometry.

    Results

    Extensive experiments demonstrate that the method generates highly realistic dual-semantic 3D illusions in only 3–5 minutes, significantly outperforming existing methods in terms of:

  • Geometric integrity
  • Semantic recognizability
  • Efficiency
---

*Auto-collected on 2026-06-20*

Tags

#3d-visual-illusion#text-to-3d#computer-vision#diffusion-models#training-free#clip#sdf#generative-ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981546