首页|期刊导航|北京交通大学学报|基于聚类算法的高铁旅客市场细分方法研究

基于聚类算法的高铁旅客市场细分方法研究OA

Research on market segmentation method for high-speed rail passengers based on clustering algorithm

中文摘要英文摘要

针对传统聚类算法进行高速铁路旅客市场细分结果不稳定,难以有效提取各类别旅客特征差异等问题,提出基于Halton序列的K-均值-自适应学习粒子群(K-Means-Adaptive Learning Par-ticle Swarm Optimization,KM-ALPSO)算法.该算法基于京沪高铁客票数据研究高铁旅客市场细分.首先,将旅客年龄、提前购票时间、出发时间段作为特征变量,利用近邻传播(Affinity Propaga-tion,AP)算法识别数据集中的样本代表点;然后,基于 Halton序列生成初始粒子群,使用 KM-ALPSO算法对样本代表点进行聚类,并与经典的K-均值算法、K-均值-粒子群(K-Means-Particle Swarm Optimization,KM-PSO)算法等进行对比分析,选择轮廓系数(Silhouette Coefficient,SC)、Davies-Bouldin(DB)、Calinski-Harabasz(CH)指标评价算法的聚类效果,确定最终聚类数目;最后,分析不同类别旅客群体的特征差异,应用频繁模式增长(Frequent Pattern-Growth,FP-Growth)算法提取强关联规则.研究结果表明:使用AP算法预处理可使基于Halton序列的KM-ALPSO算法的运行时间缩短至原来的23.26%,同时保持与未预处理时相近的评价指标值;利用Halton序列初始化粒子群位置,可以提升KM-ALPSO算法全局搜索能力;基于Halton序列的KM-ALPSO算法在客票数据分析中的聚类评价指标SC为0.332,DB为0.933,CH为708.5,表明聚类数目为5时性能最佳,且优于其他算法;5类旅客在年龄、提前购票时间和出发时间段上差异显著,表现出不同的出行计划性和出发时间偏好.

To address the instability and difficulty in effectively extracting characteristic differences among passenger groups when applying traditional clustering algorithms to high-speed rail passenger market segmentation,this study proposes a K-Means-Adaptive Learning Particle Swarm Optimiza-tion(KM-ALPSO)algorithm based on the Halton sequence.This algorithm is applied to segment the high-speed rail passenger market using ticket data from the Beijing-Shanghai high-speed railway.First,considering passenger age,advance booking time,and departure period as feature variables,the Affinity Propagation(AP)algorithm is adopted to identify representative sample points within the dataset.Second,an initial particle swarm is generated based on the Halton sequence,and the KM-ALPSO algorithm is used to cluster the representative sample points.A comparative analysis is then conducted against the classic K-Means algorithm and the K-Means-Particle Swarm Optimization(KM-PSO)algorithm.The Silhouette Coefficient(SC),Davies-Bouldin(DB)index,and Calinski-Harabasz(CH)index are selected to evaluate the clustering performance and determine the optimal number of clusters.Finally,the characteristic differences among various passenger groups are ana-lyzed,and the Frequent Pattern-Growth(FP-Growth)algorithm is employed to extract strong associa-tion rules.The results indicate that preprocessing with the AP algorithm reduces the runtime of the Halton sequence-based KM-ALPSO algorithm to 23.26%of its original duration while maintaining evaluation metrics comparable to those obtained without preprocessing.Furthermore,initializing par-ticle swarm positions with the Halton sequence enhances the global search capability of the KM-ALPSO algorithm.In the analysis of ticket data,the Halton sequence-based KM-ALPSO algorithm achieves an SC of 0.332,a DB of 0.933,and a CH of 708.5,demonstrating that a cluster number of five yields optimal performance and outperforms the baseline algorithms.The five identified passenger groups exhibit significant differences in age,advance booking time,and departure period,revealing distinct travel planning tendencies and departure time preferences.

范家乐;景云;徐彦

北京交通大学 交通运输学院,北京 100044北京交通大学 交通运输学院,北京 100044中国国家铁路集团有限公司 客运中心,北京 100844

交通工程

铁路运输市场细分聚类算法高铁旅客强关联规则

railway transportationmarket segmentationclustering algorithmhigh-speed rail passen-gersstrong association rules

《北京交通大学学报》 2026 (3)

63-72,10

国家自然科学基金(52372300) National Natural Science Foundation of China(52372300)

10.11860/j.issn.1673-0291.20250139

评论