DeduBB: Binary Code Size Reduction via Post-Link Basic Block Deduplication


Binary sizes of upgraded versions of software applications tend to be larger, primarily due to feature bloat. This poses various challenges, particularly for mobile applications. It affects upgrade rates directly impacting revenues, increases maintenance costs of supporting multiple versions, and prevents some users from getting critical security fixes. Code bloat also poses a problem for large warehouse-scale applications. Such applications experience performance degradation when their code size exceeds what smaller and more efficient code models can handle.
In this paper, we introduce a post-link optimization technique called DeduBB, which deduplicates basic blocks of an application across procedure boundaries. As the prior techniques used function outlining to deduplicate identical code sequences, they missed out on many opportunities such as duplicate code patterns that manipulate the program stack. In addition, previous techniques were either limited to the scope of a module or lacked scalable implementations required to handle large warehouse-scale applications. Our technique, DeduBB, exploits inter-module opportunities and de-duplicates more code patterns than prior techniques as it uses a novel save-and-jump code sequence to execute deduplicated code blocks. In addition, DeduBB has been designed to work on scalable post-link optimizers and can even be applied to large warehouse-scale data center applications. Finally, DeduBB is profile-guided and can be applied selectively to infrequently executed cold basic blocks to not affect application performance. In fact, in several cases, the performance of the smaller application binary improves slightly due to reductions in its hot working set size. We have designed our technique for the state-of-the-art post-link optimizers, BOLT and Propeller. Experiments show that we can significantly reduce the code size of several benchmarks by 1.55% to 18.63%, on both Arm and x86 platforms, even on binaries that have already been heavily optimized for size using existing code size reduction features. For warehouse-scale binaries, DeduBB reduces code size by up to 25.8%. Finally, aided by profiles, our technique can retain over 82% of the maximal code size savings without affecting performance.
Mon 15 JunDisplayed time zone: Mountain Time (US & Canada) change
15:50 - 17:10 | Session 2: Binary Optimization & System SecurityLCTES at Flatirons 3 Chair(s): Prasad Kulkarni University of Kansas | ||
15:50 22mTalk | DeduBB: Binary Code Size Reduction via Post-Link Basic Block Deduplication LCTES Chaitanya Mamatha Ananda University of California Riverside, Mahbod Afarin University of California, Riverside, Rajiv Gupta University of California at Riverside, Sriraman Tallam Google Inc., Han Shen Google Inc, Xinliang Li Google DOI | ||
16:12 22mTalk | SymFlow: Event-Chain-Aware Symbolic Execution for Serverless Sensitive Data Flow Detection LCTES Yuanpeng Wang Peking University, Zhineng Zhong Key Laboratory of High-Confidence Software Technologies (MOE), School of Computer Science, Peking University, Zhenkai Liang National University of Singapore, Ding Li Peking University, Yao Guo Peking University, Xiangqun Chen Peking University DOI | ||
16:34 10mShort-paper | CVS: A Metric for Security-Aware Compilation against Side-Channel Attacks in Edge SoCs (WIP)RecordedRemote LCTES Yi Han College of Computer Science and Technology, National University of Defense Technology, Changsha, China & Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China, Puhong Lei Hunan Greatwall Galaxy Science and Technology Co.,Ltd Changsha, P.R. China, Yang Shi National University of Defense Technology, Zhe Li College of Computer Science and Technology, National University of Defense Technology, Changsha, China & Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China, Xing Mou College of Computer Science and Technology, National University of Defense Technology, Changsha, China & Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China, Jianjun Chen College of Computer Science and Technology, National University of Defense Technology, Changsha, China & Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China, Yaohua Wang College of Computer Science and Technology, National University of Defense Technology, Changsha, China & Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China DOI | ||
16:44 10mShort-paper | A Programming Model for Efficient Inter-Kernel Control-Flow on Memory-Mapped Near-Data Processing Architecture (WIP) LCTES Seungheon Lee POSTECH, Wonhyuk Yang POSTECH, Seonyeong Heo Kyung Hee University, Gwangsun Kim POSTECH / Arm DOI | ||
16:54 10mShort-paper | FLUX: Frequency Scaling with Layer-wise Utilization for Energy-Efficient NPU Execution (WIP) LCTES Inho Lee Hanyang University, Ky Yeop Lim , Hyejun Kim Yonsei University, Beomseok Kim Seoul National University, Dongsuk Jeon Seoul National University, Hunjun Lee Hanyang University, Yongjun Park Yonsei University DOI | ||