基于广域分布式推理网络的PD分离架构与实现OA
PD separation architecture and implementation based on wide-area distributed inference network
随着大模型推理需求的爆发式增长,传统集中式或静态多数据中心部署模式在时延、数据合规性与资源弹性方面面临严峻挑战.基于此,提出一种基于云边协同的广域分布式推理网络架构,侧重于构建面向算力互联网的新型智算服务体系.该架构引入预填充和解码(Prefill-Decode,PD)分离机制,将低时延敏感的预填充阶段下沉至靠近数据源的边缘节点,而高吞吐的解码阶段部署于中心云,通过广域网实现安全协同.
With the explosive growth in demand for large models inference,traditional centralized or static multi-data center deployment models face severe challenges in latency,data compliance,and resource elasticity.This paper proposes a cloud-edge collaborative wide-area distributed inference network architecture,focusing on building a new intelligent-computing service system for the emerging computing-power internet.The architecture introduces a prefill-decode separation mechanism:the latency-sensitive prefill stage is offloaded to edge nodes closer to data sources,while the high-throughput decode stage is deployed in the central cloud,enabling secure collaboration over a wide-area network.
王飞飞;邓桓;唐静;王巍;苏越
中国电信股份有限公司研究院,北京 102209中国电信股份有限公司研究院,北京 102209中国电信股份有限公司研究院,北京 102209中国电信股份有限公司研究院,北京 102209中国信息通信研究院云计算与数字化研究所,北京 100191
信息技术与安全科学
广域分布式推理预填充和解码分离大模型推理
wide-area distributed inferenceprefill-decode separationlarge model inference
《信息通信技术与政策》 2026 (2)
18-23,6
评论