English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

HoYoverse Co-founder's Anuttacon Unveils LPM 1.0: A Breakthrough in Video Character Performance Generation

Forum topic · ✨步子哥 · 2026-04-23

Summary

LPM 1.0 (Large Performance Model) is a video character performance generation model from Anuttacon, the AI company founded by miHoYo co-founder Cai Haoyu, with its research paper released on arXiv in April 2026. The model converts static images into digital characters capable of real-time conversation with subtle micro-expressions and body language, targeting the 'performance trilemma' of expressiveness, real-time inference, and long-duration stability. Built on a 17-billion-parameter Diffusion Transformer, it supports full-duplex real-time dialogue (simultaneously processing user speech and AI character speech), unlimited-length streaming generation with strong identity consistency (demonstrated in 45+ minute continuous videos), and multimodal control via image/reference video, audio, and text prompts, with zero-shot generalization across realistic, 2D anime, 3D game, and non-human character styles. The team also proposed LPM-Bench, where LPM 1.0 reportedly achieves SOTA results. Currently research-only with no source code, API, or commercial release, the model is seen as a foundation for AI-driven game NPCs, virtual streamers, and conversational agents in miHoYo's long-term virtual world ambitions.

Anuttacon, the AI company founded by miHoYo co-founder Cai Haoyu in Singapore, released the paper for LPM 1.0 (Large Performance Model) on arXiv in April 2026. The model turns static images into digital characters capable of real-time conversation with fine-grained micro-expressions and body movements, achieving high identity consistency and long-duration stability. It is viewed as another step toward miHoYo's vision of "building a virtual world for a billion people by 2030."

Key Highlights

LPM 1.0 is designed for high-consistency video character performance generation, addressing the "performance trilemma" where traditional video models struggle to balance expressiveness, real-time inference, and long-duration stability:

  • Full-duplex real-time dialogue: The model simultaneously processes two audio streams — the user speaking (driving the character's listening reactions) and the AI character speaking (driving lip-sync) — enabling low-latency streaming inference and unlimited-duration continuous interaction. Official demos show videos playing continuously for over 45 minutes while the character's appearance and identity remain stable.
  • Unlimited duration with highly stable identity: Traditional models suffer from character feature drift or collapse over long generations. LPM 1.0's online streaming architecture preserves identity consistency even across hours of continuous generation, with delicate details in micro-expressions, gaze, and body rhythm.
  • Multimodal control: The model accepts image/reference video + audio + text prompts as input, supporting zero-shot generalization to realistic, 2D anime, 3D game styles, and even non-humanoid characters without per-character fine-tuning. Text controls actions, audio drives emotional expression, and images define character identity, enabling director-level control.
  • Application scenarios: LPM 1.0 is positioned as a visual engine for conversational agents, virtual streaming, and game NPCs, turning a single image into a digital human that can speak, listen, and react in real time.
  • Technical Architecture and Training

    LPM 1.0 uses a 17-billion-parameter Diffusion Transformer (DiT) architecture with multimodal conditioning. The team built a human-centric multimodal dataset with strict curation of talking-listening audio-video pairs, performance understanding, and identity-aware reference extraction. Training proceeds in two stages: first a Base LPM (17B bidirectional DiT), then distillation into an Online LPM (causal streaming generator) for low-latency, unlimited-length real-time interaction. The team also proposed LPM-Bench, a benchmark for evaluating interactive character performance, on which LPM 1.0 reportedly achieves SOTA across all evaluation dimensions.

    Background and miHoYo Connection

    Anuttacon focuses on interactive content and AGI products. It previously released the anime-style chat model "AnuNeko" and the AI-driven game *Whispers from the Star*. LPM 1.0 reflects Cai Haoyu's continued investment in the fusion of AI and gaming. The model is currently research-only — no source code, API, or commercial availability.

    Community Reactions

  • Positive: Commenters praised its long-duration consistency and expressive subtlety, with some describing it as the most emotionally convincing among comparable video models.
  • Skeptical/neutral: Others noted it focuses narrowly on character performance rather than breadth, joked about miHoYo's "digital waifu"路线 tendencies (dubbed the "digital companion" route), and cautioned that it remains at the paper stage with no usable product.

Outlook

LPM 1.0 is a focused, pragmatic step in the miHoYo/Anuttacon AI strategy — targeting "character performance," the domain miHoYo knows best, rather than pursuing general-purpose scale. If the technology lands in titles like *Genshin Impact* or *Honkai*, player interaction could improve substantially. For now, demos are impressive but large-scale commercialization remains distant; follow the arXiv paper and project page for updates.

Tags

#lpm-1-0#anuttacon#mihoyo#video-generation#diffusion-transformer#digital-humans#ai-agents#game-npcs

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618667