Summary
PipeSD is a cloud-edge collaborative inference framework accepted at ICML 2026 that extends speculative decoding beyond a single machine. Instead of running large language model inference fully on-device (privacy-preserving but compute-limited) or fully in the cloud (powerful but high-latency), PipeSD takes a hybrid approach: a small model on the edge device generates draft tokens, while a large model in the cloud verifies them in batches. The system introduces two key innovations: a dynamically optimized token-batch pipeline schedule that overlaps generation and communication so neither waits on the other, and a dual-threshold Bayesian auto-tuning mechanism that decides when verification is triggered. Experiments show 1.16x-2.16x speedups and 14.3%-25.3% energy reduction. The paper's core insight is that scaling speculative decoding to cloud-edge settings requires managing the combined latency of communication and computation, making pipeline scheduling the central challenge.
Running large language model inference on edge devices presents a trade-off: go fully local (privacy-preserving, but compute-limited), or go fully cloud (ample compute, but high latency). PipeSD takes the middle path — accepted at ICML 2026 — where a small model on the edge device writes drafts and a large model in the cloud verifies them in batches.
Two Key Innovations
1. Dynamic programming-optimized token-batch pipeline scheduling — overlapping generation and communication so neither side waits on the other.
2. Dual-threshold Bayesian auto-tuned verification trigger — adaptively deciding when to trigger cloud verification.
Results
- 1.16x–2.16x speedup
- 14.3%–25.3% reduction in energy consumption
> Extending speculative decoding from a single machine to cloud-edge collaboration means the core difficulty is the combined latency of communication and computation. Pipeline scheduling is the key to solving it.
Paper Information
- Title: PipeSD: Cloud-Edge Collaborative Pipeline Inference with Speculative Decoding
- Authors: Yunhe Han, Yunqi Gao, et al.
- Venue: ICML 2026
- Preprint: arXiv:2605.13319 (cs.DC)
- Link: https://arxiv.org/abs/2605.13319
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177620153