Anuttacon, the AI company founded by miHoYo co-founder Cai Haoyu in Singapore, released the paper for LPM 1.0 (Large Performance Model) on arXiv in April 2026. The model turns static images into digital characters capable of real-time conversation with fine-grained micro-expressions and body movements, achieving high identity consistency and long-duration stability. It is viewed as another step toward miHoYo's vision of "building a virtual world for a billion people by 2030."
Key Highlights
LPM 1.0 is designed for high-consistency video character performance generation, addressing the "performance trilemma" where traditional video models struggle to balance expressiveness, real-time inference, and long-duration stability:
- Full-duplex real-time dialogue: The model simultaneously processes two audio streams — the user speaking (driving the character's listening reactions) and the AI character speaking (driving lip-sync) — enabling low-latency streaming inference and unlimited-duration continuous interaction. Official demos show videos playing continuously for over 45 minutes while the character's appearance and identity remain stable.
- Unlimited duration with highly stable identity: Traditional models suffer from character feature drift or collapse over long generations. LPM 1.0's online streaming architecture preserves identity consistency even across hours of continuous generation, with delicate details in micro-expressions, gaze, and body rhythm.
- Multimodal control: The model accepts image/reference video + audio + text prompts as input, supporting zero-shot generalization to realistic, 2D anime, 3D game styles, and even non-humanoid characters without per-character fine-tuning. Text controls actions, audio drives emotional expression, and images define character identity, enabling director-level control.
- Application scenarios: LPM 1.0 is positioned as a visual engine for conversational agents, virtual streaming, and game NPCs, turning a single image into a digital human that can speak, listen, and react in real time.
- Positive: Commenters praised its long-duration consistency and expressive subtlety, with some describing it as the most emotionally convincing among comparable video models.
- Skeptical/neutral: Others noted it focuses narrowly on character performance rather than breadth, joked about miHoYo's "digital waifu"路线 tendencies (dubbed the "digital companion" route), and cautioned that it remains at the paper stage with no usable product.
Technical Architecture and Training
LPM 1.0 uses a 17-billion-parameter Diffusion Transformer (DiT) architecture with multimodal conditioning. The team built a human-centric multimodal dataset with strict curation of talking-listening audio-video pairs, performance understanding, and identity-aware reference extraction. Training proceeds in two stages: first a Base LPM (17B bidirectional DiT), then distillation into an Online LPM (causal streaming generator) for low-latency, unlimited-length real-time interaction. The team also proposed LPM-Bench, a benchmark for evaluating interactive character performance, on which LPM 1.0 reportedly achieves SOTA across all evaluation dimensions.
Background and miHoYo Connection
Anuttacon focuses on interactive content and AGI products. It previously released the anime-style chat model "AnuNeko" and the AI-driven game *Whispers from the Star*. LPM 1.0 reflects Cai Haoyu's continued investment in the fusion of AI and gaming. The model is currently research-only — no source code, API, or commercial availability.
Community Reactions
Outlook
LPM 1.0 is a focused, pragmatic step in the miHoYo/Anuttacon AI strategy — targeting "character performance," the domain miHoYo knows best, rather than pursuing general-purpose scale. If the technology lands in titles like *Genshin Impact* or *Honkai*, player interaction could improve substantially. For now, demos are impressive but large-scale commercialization remains distant; follow the arXiv paper and project page for updates.