首页|期刊导航|硅酸盐学报|低碳熟料强度预测的数据增强建模及模型可解释性分析

低碳熟料强度预测的数据增强建模及模型可解释性分析OA

Data-Augmented Modeling and Model Interpretability Analysis for Strength Prediction of Low-Carbon Clinker

中文摘要英文摘要

针对低碳模拟熟料体系样本量小、组成空间稀疏,导致代理模型跨体系预测能力不足的问题,本工作提出一种自迭代数据增强框架,用于优化其28 d抗压强度预测模型.该框架在组成约束条件下生成虚拟组成,并由上一轮模型赋予伪标签以逐轮自训练.采用测试集全程隔离、验证集仅包含实验数据的锚定评估策略,抑制由伪标签引入造成的误差传递.结果表明:迭代增强后的代理模型测试误差降低,跨体系泛化能力提升.在第2轮迭代中,以模式B采样训练得到的代理模型在体系内拟合与体系外泛化之间取得较优平衡,其总体均方误差(MSE)为 9.8,体系内、外 MSE 分别为 6.9、11.3,相关系数 R2为0.79,平均绝对误差与平均绝对百分比误差分别为2.36与5.11%.在此基础上,进一步结合受约束的置换特征重要性、累积局部效应与沙普利加性解释开展分析,识别关键矿相的全局敏感性、平均响应关系与样本级贡献,为模型可靠性审查、关键矿相筛选及后续模拟熟料组成优化提供可检验的参考依据.

Introduction The clinker calcination process involves high-temperature solid/liquid reactions,and the controllable preparation of multiple clinker phase compositions,which further affect subsequent synergistic hydration reactions.The relationship between composition and 28 d compressive strength is not simply linear,but rather a complex mapping governed by the coupling of multiple processes.Conventional point-by-point experimental screening makes it difficult to cover a sufficiently broad compositional space within a limited timeframe.This restricts the exploration efficiency of novel low-carbon clinker.Machine learning models can be used to learn the composition and property relationships between clinker phase and strength,and to support the screening of candidate clinker compositions.However,under small data conditions,these models are susceptible to insufficient training distribution coverage,resulting in a unstable cross-system predictive performance.Therefore,this study was to develop a self-iterative data augmentation framework to address the above-mentioned issues.This framework was to expand effective supervised information to optimize the model's predictive performance in sparsely sampled compositional regions.Subsequently,model interpretability analysis was conducted to identify the key clinker phases affecting 28 d compressive strength prediction and their response patterns.This could provide a testable model basis for the optimization of low-carbon clinker compositions. Methods This study defined each simulated clinker as one sample.The dataset contained 264 simulated clinker samples,with the compressive strengths ranging from 0 to 70.3 MPa.The model inputs consisted of 12 features,including 11 clinker phases and gypsum,while the output was the 28 d compressive strength.In the self-iterative data augmentation strategy,an artificial neural network(ANN)model was first trained using the experimental data.Virtual data were then generated under compositional constraints,and their 28 d compressive strengths were predicted by the model from the previous iteration to assign pseudo-labels.The virtual and experimental samples were subsequently combined for training,and the surrogate model was updated iteratively.Two sampling patterns were adopted for virtual sample generation,i.e.,1)Pattern A:dense sampling within the experimental samples'compositional systems,and 2)Pattern B:generating virtual samples at a ratio of 1:1 from the in-and out-of-system regions of the experimental samples'compositional space.Before training,36 experimental samples were set aside as a test set,which remained isolated throughout training,iteration and hyperparameter selection.The validation set consisted exclusively of real experimental samples,whereas virtual samples were included only in the training set.Model hyperparameters were determined based on the Bayesian optimization,with the validation-set mean squared error(MSE)as the selection criterion.Finally,the model predictive performance was evaluated by MSE,the coefficient of determination(R2),mean absolute error(MAE),and mean absolute percentage error(MAPE).After the final surrogate model was determined,multilevel model interpretation was further performed via combining constrained permutation feature importance,accumulated local effects and SHapley additive explanations(SHAP). Results and discussion The self-iterative augmentation strategy effectively fills sparsely sampled regions,while maintaining feasible-domain constraints.This provides richer training information for subsequent accuracy evaluation and interpretability analysis of the surrogate model.Meanwhile,the data augmentation process does not generate a large number of near-duplicate samples,and the training set retains the necessary heterogeneity.The inclusion of virtual samples reduces the model's sensitivity to randomness in data partitioning.Moreover,the overall test set error and the in-/out-of-system errors decrease progressively over successive iterations,indicating an improved cross-system generalization performance.Pattern B in Iteration II achieves a favorable balance between predictive accuracy and extrapolation stability.Its overall test set MSE is 9.8,with in-system and out-of-system MSE values of 6.9 and 11.3,respectively.The R2,MAE,and MAPE are 0.79,2.36 and 5.11%,respectively.These results indicate that,under the current small data conditions,data augmentation can substantially improve the model's ability to identify performance trends.Subsequent interpretability analysis of the optimal surrogate model shows that different clinker phases can be classified as positive or negative driving phases.The analysis further indicates that when C3A content exceeds approximately 10%,its positive driving effect on strength decreases rapidly.Although CA and C12A7 act as positive driving phases,their benefits diminish when their contents exceed 40%,and even shift to strongly negative responses. Conclusions The self-iterative data augmentation framework proposed could expand a supervised training information for low-carbon clinker systems characterized by small datasets and sparse compositional spaces.The evaluation bias introduced by pseudo-labels could be also suppressed under a rigorous evaluation strategy.The results showed that data augmentation significantly improved the prediction accuracy and cross-system generalization capability of the ANN surrogate model for 28 d compressive strength.In Iteration II,Pattern B achieved the optimum overall performance under the current data conditions.Although further iterations reduced some in-system errors,an error rebound was observed for out-of-system samples.This indicated that diminishing marginal information could gain from pseudo-labels constrained the model's generalization performance.The combined use of constrained PFI,ALE,and SHAP revealed the decision-making basis of the surrogate model in terms of global sensitivity,average response trends,and sample-level contributions.The results indicated that C4A3$ and CA were the main positive driving clinker phases for compressive strength enhancement,whereas C3A,C5S2$ and C2S were the main negative factors.Moreover,C12A7 exhibited a pronounced threshold effect.These findings could provide a testable reference for subsequent composition screening,constraint-boundary definition and inverse design of low-carbon clinker.

翟牧楠;郅晓;叶家元;任雪红;吴春丽;张洪滔;崔文娟;颜景华;张文生

中国建筑材料科学研究总院有限公司,北京 100024中国建筑材料科学研究总院有限公司,北京 100024中国建筑材料科学研究总院有限公司,北京 100024中国建筑材料科学研究总院有限公司,北京 100024中存大数据科技有限公司,北京 100024中国建筑材料科学研究总院有限公司,北京 100024中国建筑材料科学研究总院有限公司,北京 100024中国建筑材料科学研究总院有限公司,北京 100024中国建筑材料科学研究总院有限公司,北京 100024

化学化工

低碳熟料28 d抗压强度预测模型自迭代数据增强可解释性

low-carbon clinker28 d compressive strengthprediction modelself-iterativedata augmentationmodel interpretability

《硅酸盐学报》 2026 (8)

2627-2643,17

"十四五"国家重点研发计划(2022YFC3803101)国家自然科学基金双碳专项项目(52341202).

10.14062/j.issn.0454-5648.20260127

评论