基于随机森林算法和数据增强策略双驱动的多尺度疲劳裂纹扩展预测研究

  • 刁圣轩 ,
  • 陈永保 ,
  • 肖金雍 ,
  • 单康中 ,
  • 杨杰
展开
  • (上海理工大学 能源与动力工程学院 上海市动力工程多相流动与传热重点实验室  上海 200093)

收稿日期: 2024-08-28

  修回日期: 2025-03-10

  网络出版日期: 2025-04-08

基金资助

国家自然科学基金项目;国家自然科学基金项目

Multi-Scale Fatigue Crack Propagation Prediction Based on the Dual Drive of Random Forest Algorithm and Data Augmentation Strategy

  • DIAO Ku-Han ,
  • CHEN Yong-Bao ,
  • XIAO Jin-Yong ,
  • SHAN Kang-Zhong ,
  • YANG Jie
Expand
  • Shanghai Key Laboratory of Multiphase Flow and Heat Transfer in Power Engineering, School of Energy and Power Engineering, University of Shanghai for Science and Technology, Shanghai 200093, China

Received date: 2024-08-28

  Revised date: 2025-03-10

  Online published: 2025-04-08

摘要

为了实现多尺度疲劳裂纹扩展过程的准确预测,并增强预测模型的可解释性和泛化能力,本工作选用304奥氏体不锈钢为研究对象,首先对K近邻回归(KNN)、支持向量机回归(SVR)和随机森林回归(RF) 3种算法的疲劳裂纹扩展预测能力进行了对比,筛选出预测性能最优的算法。随后,选择最优算法,分别基于长裂纹、短裂纹和多尺度3种疲劳裂纹扩展模型添加数据增强策略进行数据增强,以提高疲劳裂纹扩展预测的准确性。最后,基于添加数据增强策略的双驱动模型框架,进行不同载荷下的多尺度疲劳裂纹扩展预测。结果表明,与SVR和KNN算法相比,RF算法更适合于疲劳裂纹扩展预测,其基于多尺度疲劳裂纹扩展模型的数据增强效果明显优于其余2种模型。基于RF算法和数据增强策略的双驱动模型具有很好的泛化能力,且相比基于RF算法纯数据驱动模型,双驱动模型预测结果更加准确。

本文引用格式

刁圣轩 , 陈永保 , 肖金雍 , 单康中 , 杨杰 . 基于随机森林算法和数据增强策略双驱动的多尺度疲劳裂纹扩展预测研究[J]. 金属学报, 0 : 0 -0 . DOI: 10.11900/0412.1961.2024.00302

Abstract

Data-driven methods based on machine learning have been employed to predict fatigue crack propagation. However, existing studies have largely overlooked the multi-scale nature of this process. Relying solely on macroscopic data for long-crack prediction often fails to capture the complete crack growth process, potentially resulting in non-conservative predictions. Moreover, purely data-driven models often lack interpretability, exhibit limited generalization capabilities, and struggle to adhere to physical laws. The challenge of integrating modeling with reasonable interpretations, driven by both data and mechanisms, remains a significant issue for researchers. In this study, we selected 304 austenitic stainless steel as the research object. To identify the algorithm with the best predictive performance, firstly, the fatigue crack propagation prediction capabilities of three algorithms were compared: K-nearest neighbor regression (KNN), support vector machine regression (SVR), and random forest regression (RF). The most effective algorithm and implemented data augmentation strategies based on three fatigue crack propagation models were selected, namely, long cracks, short cracks, and multi-scale, to enhance prediction accuracy. Finally, using a dual-drive model framework that incorporated the data augmentation strategy, predictions of multi-scale fatigue crack propagation under different loads were conducted. Compared to the SVR and KNN algorithms, the results indicated that the RF algorithm had a lower RMSE value and higher coefficient of determination (R2), making it more suitable for predicting fatigue crack propagation. Under a load of 370 MPa, the prediction accuracy for the training set ranked in the order of RF > KNN > SVR. By contrast, the accuracy for the test set was in the order of RF > SVR > KNN. The multi-scale fatigue crack propagation model effectively captured the entire process of crack growth, whereas the long- and short-crack models accurately represented only parts of it. Data enhancement based on the multi-scale model demonstrated significantly better results than the other two models, with increases in accuracy of 25.76% and 71.74% for the training and test sets, respectively. The dual-drive model based on the RF algorithm and data enhancement strategy exhibited strong generalization capabilities. Under a load of 350 MPa, the RMSE values for the training and test sets were 0.046 and 0.111, respectively, and R2 reached 0.995 and 0.961. Under a load of 330 MPa, the RMSE values for the training and test sets were 0.171 and 0.081, respectively, and the R2 reached 0.911 and 0.721. Finally, compared with the pure data-driven model based on the RF algorithm, the predictions from the dual-drive model were found to be significantly more accurate.

文章导航

/