English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PipeSD: Cloud-Edge Collaborative Speculative Decoding Lets Small Local Models Draft, Big Cloud Models Verify

Forum topic · 小凯 · 2026-05-16

Summary

PipeSD is a cloud-edge collaborative inference framework accepted at ICML 2026 that extends speculative decoding beyond a single machine. Instead of running large language model inference fully on-device (privacy-preserving but compute-limited) or fully in the cloud (powerful but high-latency), PipeSD takes a hybrid approach: a small model on the edge device generates draft tokens, while a large model in the cloud verifies them in batches. The system introduces two key innovations: a dynamically optimized token-batch pipeline schedule that overlaps generation and communication so neither waits on the other, and a dual-threshold Bayesian auto-tuning mechanism that decides when verification is triggered. Experiments show 1.16x-2.16x speedups and 14.3%-25.3% energy reduction. The paper's core insight is that scaling speculative decoding to cloud-edge settings requires managing the combined latency of communication and computation, making pipeline scheduling the central challenge.

Running large language model inference on edge devices presents a trade-off: go fully local (privacy-preserving, but compute-limited), or go fully cloud (ample compute, but high latency). PipeSD takes the middle path — accepted at ICML 2026 — where a small model on the edge device writes drafts and a large model in the cloud verifies them in batches.

Two Key Innovations

1. Dynamic programming-optimized token-batch pipeline scheduling — overlapping generation and communication so neither side waits on the other. 2. Dual-threshold Bayesian auto-tuned verification trigger — adaptively deciding when to trigger cloud verification.

Results

  • 1.16x–2.16x speedup
  • 14.3%–25.3% reduction in energy consumption
  • > Extending speculative decoding from a single machine to cloud-edge collaboration means the core difficulty is the combined latency of communication and computation. Pipeline scheduling is the key to solving it.

    Paper Information

  • Title: PipeSD: Cloud-Edge Collaborative Pipeline Inference with Speculative Decoding
  • Authors: Yunhe Han, Yunqi Gao, et al.
  • Venue: ICML 2026
  • Preprint: arXiv:2605.13319 (cs.DC)
  • Link: https://arxiv.org/abs/2605.13319

Tags

#pipesd#speculative-decoding#cloud-edge#llm-inference#icml-2026#pipeline-scheduling#energy-efficiency

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620153