FLUX: Frequency Scaling with Layer-wise Utilization for Energy-Efficient NPU Execution (WIP)
With the widespread adoption of Deep Neural Networks (DNNs), Neural Processing Units (NPUs) are emerging as energy-efficient alternatives to GPUs through parallel processing and high data reuse.
However, since diverse deep learning kernels have different memory and computation resource requirements, a utilization imbalance between memory and computation resources often occurs.
To address this challenge, we propose FLUX, a frequency-scalable NPU system that applies Dynamic Frequency Scaling (DFS) to the core's frequency based on layer-wise arithmetic intensity.
FLUX splits the clocking into an adjustable core domain and a fixed system domain.
The core-domain frequency is first estimated via roofline cycle analysis and then refined at runtime by an Energy-Delay Product (EDP)-driven calibration.
Evaluation on a Gemmini NPU using a 28nm process technology shows that FLUX achieves 5.7%, 16.5%, and 27.9% EDP improvements on 8×8, 16×16, and 32×32 systolic arrays for ResNet50, respectively, reaching within 1.7% of the oracle optimal frequency assignment.
Mon 15 JunDisplayed time zone: Mountain Time (US & Canada) change
15:50 - 17:10 | Session 2: Binary Optimization & System SecurityLCTES at Flatirons 3 Chair(s): Prasad Kulkarni University of Kansas | ||
15:50 22mTalk | DeduBB: Binary Code Size Reduction via Post-Link Basic Block Deduplication LCTES Chaitanya Mamatha Ananda University of California Riverside, Mahbod Afarin University of California, Riverside, Rajiv Gupta University of California at Riverside, Sriraman Tallam Google Inc., Han Shen Google Inc, Xinliang Li Google DOI | ||
16:12 22mTalk | SymFlow: Event-Chain-Aware Symbolic Execution for Serverless Sensitive Data Flow Detection LCTES Yuanpeng Wang Peking University, Zhineng Zhong Key Laboratory of High-Confidence Software Technologies (MOE), School of Computer Science, Peking University, Zhenkai Liang National University of Singapore, Ding Li Peking University, Yao Guo Peking University, Xiangqun Chen Peking University DOI | ||
16:34 10mShort-paper | CVS: A Metric for Security-Aware Compilation against Side-Channel Attacks in Edge SoCs (WIP)RecordedRemote LCTES Yi Han College of Computer Science and Technology, National University of Defense Technology, Changsha, China & Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China, Puhong Lei Hunan Greatwall Galaxy Science and Technology Co.,Ltd Changsha, P.R. China, Yang Shi National University of Defense Technology, Zhe Li College of Computer Science and Technology, National University of Defense Technology, Changsha, China & Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China, Xing Mou College of Computer Science and Technology, National University of Defense Technology, Changsha, China & Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China, Jianjun Chen College of Computer Science and Technology, National University of Defense Technology, Changsha, China & Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China, Yaohua Wang College of Computer Science and Technology, National University of Defense Technology, Changsha, China & Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China DOI | ||
16:44 10mShort-paper | A Programming Model for Efficient Inter-Kernel Control-Flow on Memory-Mapped Near-Data Processing Architecture (WIP) LCTES Seungheon Lee POSTECH, Wonhyuk Yang POSTECH, Seonyeong Heo Kyung Hee University, Gwangsun Kim POSTECH / Arm DOI | ||
16:54 10mShort-paper | FLUX: Frequency Scaling with Layer-wise Utilization for Energy-Efficient NPU Execution (WIP) LCTES Inho Lee Hanyang University, Ky Yeop Lim , Hyejun Kim Yonsei University, Beomseok Kim Seoul National University, Dongsuk Jeon Seoul National University, Hunjun Lee Hanyang University, Yongjun Park Yonsei University DOI | ||