首页|期刊导航|Security and Safety|PUMA: Secure inference of LLaMA-7B in five minutes

PUMA: Secure inference of LLaMA-7B in five minutesOACSCD

中文摘要

Transformer models(e.g., Bert and GPT) have shown their dominance in machine learning tasks. Many cloud companies have begun to provide services based on Transformer models, examples include translation and text-speech conversion. However, such a service inevitably requires access to the client''s data, which might contain sensitive information. Theoretically, running the services under secure multi-party computation(MPC)could protect clients'' privacy. However, current MPC frameworks are still limited in terms of model performance, efficiency, deployment, and functionality, especially when facing complex Transformer models. To this end, we propose an MPC framework Pum A to enable secure and efficient Transformer model inference. We first design high-quality approximations for the bottleneck functions in Transformers such as GELU and Softmax, reducing about20-76% computation and communication costs than state-of-the-art works without performance drop. Then, we provide concrete instantiations for secure Embedding and Layer Norm.These implementations produce correct results and integrate compatible system architectures of cleartext Transformer models. Finally, we conducted extensive experiments on six popular benchmarks: text classification/generation/summarization/translation, audio-to-text,and image-to-text. Results show that Pum A can finish most tasks in several minutes, with comparable model performance(e.g., accuracy) as cleartext, and even evaluate LLa MA-7B in less than 5 minutes to generate 1 token.

Ye Dong;Wen-Jie Lu;Yancheng Zheng;Haoqi Wu;Derun Zhao;Jin Tan;Zhicong Huang;Cheng Hong;Tao Wei;Wen-Guang Chen;Jianying Zhou

Ant Group,Beijing 100081,China National University of Singapore,Singapore 119260,SingaporeAnt Group,Beijing 100081,ChinaAnt Group,Beijing 100081,ChinaAnt Group,Beijing 100081,ChinaAnt Group,Beijing 100081,ChinaAnt Group,Beijing 100081,ChinaAnt Group,Beijing 100081,ChinaAnt Group,Beijing 100081,ChinaAnt Group,Beijing 100081,ChinaAnt Group,Beijing 100081,ChinaSingapore University of Technology and Design,Singapore 487372,Singapore

信息技术与安全科学

PrivacySecuritySecure Three-Party ComputationPrivacy-Preserving Machine LearningLarge Language Models

《Security and Safety》 2025 (4)

P.145-168,24

10.1051/sands/2025014

评论