Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to speculate multiple tokens, reducing expensive target model decoding steps. Its effectiveness depends heavily on draft selection, motivating adaptive methods that exploit variation across inputs and generation stages. On memory-constrained edge devices, however, these methods often fail to improve end-to-end throughput due to the overhead of switching between draft models. We identify a key limitation in this setting: the mismatch between draft selection and draft availability under tight memory budgets.

To address this challenge, we present MemSpec, a prediction-guided, memory-aware runtime for adaptive speculative decoding on edge devices. MemSpec decouples draft selection from execution through proactive resident working-set management. A lightweight predictor estimates draft effectiveness from prompt and generation context, while a memory-aware scheduler reduces reactive model loading overhead. Experiments on a Jetson Orin Nano show that MemSpec improves steady-state generation throughput by 40.7% on average over state-of-the-art bandit-based adaptive methods while closely approaching the oracle upper bound.

Tue 16 Jun

Displayed time zone: Mountain Time (US & Canada) change

15:50 - 17:00
Session 5: Memory Efficiency and Control-Flow AnalysisLCTES at Flatirons 3
Chair(s): Sandrine Blazy University of Rennes
15:50
22m
Talk
MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices
LCTES
Eunjeong Kim Kyungpook National University, Yeong Jun Jeon Kyungpook National University, Myeonggyun Han Kyungpook National University
DOI
16:12
22m
Talk
Bridging the Memory Hotness Gap in Edge Systems with Hotness-Segregated Object Allocation
LCTES
Ruizhe Huang Peking University, Jiahua Wang Peking University, Qihang Xu Peking University, Peng Jiang Southeast University, Zhida An Peking University, Ding Li Peking University, Yao Guo Peking University, Xiangqun Chen Peking University, Yuxin Ren Huawei Technologies, Ning Jia Huawei Technologies
DOI
16:34
22m
Talk
On the Origins of Indirect Jumps in Embedded SoftwareResults ReproducedArtifacts AvailableArtifacts Evaluated
LCTES
Ariane Nicolas Univ Rennes, Inria, CNRS, IRISA, Ronan Lashermes Rambus, Isabelle Puaut Université de Rennes - Inria - CNRS - IRISA, Erven Rohou Université de Rennes - Inria - CNRS - IRISA
DOI
16:56
4m
Talk
Closing Session
LCTES
Jeronimo Castrillon TU Dresden, Germany