English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DeepSlide: A Multi-Agent System for Full Presentation Preparation, From Artifact to Delivery

Forum topic · 小凯 · 2026-05-19

Summary

DeepSlide (arXiv:2505.10892) is a human-in-the-loop multi-agent system that automates the entire academic presentation workflow. Unlike most AI slide generators that focus only on producing visually plausible decks, DeepSlide optimizes the delivery process itself, including pacing, narrative flow, and rehearsal. The system combines four components: a controllable logical-chain planner with per-node time budgets, a lightweight content-tree retriever for evidence grounding, Markov-style sequential rendering with style inheritance, and sandboxed execution with minimal repair to guarantee renderability. The authors also introduce a dual-scorecard benchmark that cleanly separates static artifact quality from dynamic delivery performance. Evaluations across 20 domains and diverse audiences show DeepSlide matches strong baselines on artifact quality while achieving consistently larger gains on delivery metrics such as narrative fluency, pacing accuracy, and slide-script coordination. This forum post summarizes the paper's contributions and links to the arXiv page.

Paper Overview

  • Field: NLP
  • Authors: Ming Yang, Zhiwei Zhang, Jiahang Li
  • Published: 2025-05-15
  • arXiv: 2505.10892
  • Summary

    Presentations are a primary medium for scholarly communication, yet most AI slide generators optimize only the *artifact* — a visually plausible deck — while under-optimizing the *delivery process*: pacing, narrative, and presentation preparation.

    DeepSlide is a human-in-the-loop multi-agent system that supports preparing the full presentation process, from requirement elicitation and time-budgeted narrative planning, to evidence-grounded slide–script generation, attention augmentation, and rehearsal support.

    Key Components

    1. Controllable logical-chain planner with per-node time budgets 2. Lightweight content-tree retriever for grounding 3. Markov-style sequential rendering with style inheritance 4. Sandboxed execution with minimal repair to ensure renderability

    Evaluation

    The authors introduce a dual-scorecard benchmark that cleanly separates static artifact quality from dynamic delivery performance. Across 20 domains and diverse audiences, DeepSlide:

  • Matches strong baselines on artifact quality
  • Consistently achieves larger gains on delivery metrics, improving narrative fluency, pacing accuracy, and slide–script coordination

Original Abstract (excerpt)

> Presentations are a primary medium for scholarly communication, yet most AI slide generators optimize the artifact (a visually plausible deck) while under-optimizing the delivery process (pacing, narrative, and presentation preparation). We present DeepSlide, a human-in-the-loop multi-agent system that supports preparing the full presentation process, from requirement elicitation and time-budgeted narrative planning, to evidence-grounded slide–script generation, attention augmentation, and rehearsal support. DeepSlide integrates (i) a controllable logical-chain planner with per-node time budgets, (ii) a lightweight content-tree retriever for grounding, (iii) Markov-style sequential rendering with style inheritance, and (iv) sandboxed execution with minimal repair to ensure renderability...

---

*Auto-collected on 2026-05-19*

Tags

#deepslide#nlp#arxiv#multi-agent-systems#presentation-generation#ai#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620349