Efficient Inference for Edge Large Language Models:A SurveyOA
Large language models(LLMs)have demonstrated remarkable capabilities in natural language processing.Their massive computational and memory requirements often necessitate cloud-based deployment,introducing challenges related to cost,latency,privacy,and network reliability.Deploying ondevice LLMs alleviates these challenges,but is hindered by the severe resource constraints of edge hardware.This survey reviews efficient inference techniques for edge LLMs,with a focus on two key strategies of speculative decoding and model offloading.We categorize strategies into single-device and multi-device types,systematically analyzing the principles,recent advancements,implementations,and support within edge frameworks.Finally,we highlight the open challenges and future research directions that will advance the field of edge LLM inference.
Guanyu Cai;Ruiming Tian;Lang Yang;Yunzhe Jia;Lingkun Li;Jiliang Wang
School of Software,Tsinghua University,Beijing 100190,ChinaSchool of Software Engineering,Beijing Jiaotong University,Beijing 100044,ChinaSchool of Software,Tsinghua University,Beijing 100190,ChinaSchool of Software,Tsinghua University,Beijing 100190,ChinaSchool of Software Engineering,Beijing Jiaotong University,Beijing 100044,ChinaSchool of Software,Tsinghua University,Beijing 100190,China
信息技术与安全科学
large language models(LLMs)edge computingon-devicespeculative decodingmodel offloading
《Tsinghua Science and Technology》 2026 (3)
P.1365-1380,16
supported by the National Key R&D Program of China(No.2022YFC3801300)the National Natural Science Foundation of China(Nos.U22A2031,62172250,and 62402028)the Tsinghua University-Fuzhou Joint Institute for Data Technology.
评论