Monet: Reasoning in Latent Visual Space
A joint team from Peking University, Kuaishou, and MIT presents Monet, a framework that lets multimodal large language models (MLLMs) move beyond "describing what they see" and instead perform reasoning inside a latent visual space — a kind of "mind's eye" for AI.
Core Concept
Rather than simple pixel recognition, Monet conducts continuous mental simulation in a high-dimensional latent visual space. The approach is grounded in the manifold hypothesis: data in high-dimensional space concentrates on low-dimensional manifolds. Monet performs "mental simulation" on these low-dimensional manifolds, avoiding the curse of dimensionality.
Technical Architecture
The training pipeline consists of two components:
1. SFT (Distilled Supervised Fine-Tuning) — three stages:
- Stage 1: Warm-up adaptation to interleaved image-text reasoning
- Stage 2: Acquiring high-quality target latent embeddings
- Stage 3: Autonomously generating embeddings without auxiliary images
- Incorporates continuous latent variables into the reinforcement learning policy gradient, directly optimizing the model's "visual intuition" from reward signals.
- Disaster-response robotics: simulating complex environments to plan safe paths
- Medical prediction: simulating disease progression to assist clinical decision-making
2. VLPO (Variational Latent Policy Optimization):
Experimental Results
Monet significantly outperforms baseline models (including GPT-4V) on both standard reasoning tasks and out-of-distribution (OOD) abstract reasoning, as shown in the poster's reported accuracies:
| Model | Standard Reasoning | OOD Abstract Reasoning | |---|---|---| | Baseline (SFT + GRPO) | 48.5% | 22.0% | | Monet (SFT + VLPO) | 54.5% | 33.7% |
Future Outlook and Applications
When machines possess a "mental model," they can rehearse the consequences of actions in their minds before acting — opening a new chapter for AI in the physical world: