首页|期刊导航|信息工程大学学报|改进双延迟深度确定性策略梯度的库存控制方法

改进双延迟深度确定性策略梯度的库存控制方法OA

Inventory Control Method Based on Improved Twin Delayed Deep Deterministic Policy Gradient

中文摘要英文摘要

针对不确定需求和供应延迟环境中库存控制难度大、成本偏高的问题,提出一种改进双延迟深度确定性策略梯度(TD3)的库存控制方法.首先,将库存控制抽象为满足率最高和成本最小双目标的马尔可夫决策过程,作为TD3算法的训练环境;其次,采用优先经验回放机制,提高TD3算法的采样效率,将长短期记忆网络融入TD3算法的多层感知机,优化网络结构;最后,通过TD3算法与环境交互,实现库存控制中成本和满足率优化.实验结果表明,所提方法在达到满足率阈值条件下,其库存控制成本较原始TD3算法降低22.2%.

To address the difficulty and high cost of inventory control in the environment of uncertain demand and supply delays,an inventory control method based on the improved twin delayed deep de-terministic policy gradient(TD3)is proposed.Firstly,the inventory control is abstracted as a Markov decision process with dual objectives of service level maximization and cost minimization,which serves as the training environment for the TD3 algorithm.Secondly,the prioritized experience replay mechanism is adopted to improve the sampling efficiency of the TD3 algorithm,and the long short-term memory(LSTM)is integrated into the multi-layer perceptron of the TD3 algorithm to optimize the net-work structure.Finally,the TD3 algorithm is used to interact with the environment to optimize both cost and service level in inventory control.Experimental results demonstrate that the inventory control cost of the proposed method is 22.2%lower than that of the original TD3 algorithm when the service level threshold is reached.

龚永奇;郭基联;张亮;唐希浪

空军工程大学,陕西 西安,710038空军工程大学,陕西 西安,710038空军工程大学,陕西 西安,710038空军工程大学,陕西 西安,710038

信息技术与安全科学

库存控制双延迟深度确定性策略梯度优先经验回放长短期记忆网络

inventory controlTD3prioritized experience replayLSTM

《信息工程大学学报》 2026 (1)

35-41,7

国家自然科学基金(72201276)

10.3969/j.issn.1671-0673.2026.01.005

评论