مجله علمی  رایانش نرم و فناوری اطلاعات

مجله علمی رایانش نرم و فناوری اطلاعات

تشخیص ناسازگاری داده‌های نامتوازن در طول زمان برای پیش‌بینی نقص نرم‌افزار به موقع

نوع مقاله : مقاله پژوهشی فارسی

نویسندگان
1 دانشکده مهندسی کامپیوتر، دانشگاه خواجه نصیرالدین طوسی، تهران، ایران
2 دانشکده مهندسی کامپیوتر، دانشگاه خواجه نصیرالدین طوسی، تهران، ایران.
3 آزمایشگاه زمان‌بندی توزیع‌شده و موتور داده هوآوی، کانادا.
چکیده
چکیده- در زمینه پیش‌بینی نقص نرم‌افزار به‌هنگام (JIT-SDP)، مطالعاتی انجام شده‌اند که ناسازگاری خروجی تفسیر روی مجموعه‌داده‌ها و مدل‌های مختلف را بررسی کرده‌اند. تکنیک‌های تفسیر، اهمیت و تاثیر ویژگی‌ها را در نتیجه پیش‌بینی کلاس هدف بدون نیاز به داده‌های آزمون برچسب‌دار مشخص می‌کنند. آنها نشان داده‌اند که خروجی تفسیر یک مدل آموزش‌دیده برون‌خط روی مجموعه‌داده‌های مختلف با یکدیگر ناسازگار است. چنین ناسازگاری‌ای ممکن است در یک مجموعه‌داده وابسته به زمان نیز رخ دهد و منجر به ناپایداری مدل و در نتیجه افت عملکرد آن در طول زمان شود. در این پژوهش، یک روش جدید با عنوان تشخیص ناسازگاری تفسیر در JIT-SDP (IID-JSDP) معرفی می‌شود که با شناسایی تغییرات معنادار در تفسیر نمونه‌ها در طول زمان، ناپایداری‌های مدل را پیش‌بینی می‌کند. برای ارزیابی، از روش‌های مرجع مبتنی بر پایش نرخ خطا استفاده شده است؛ چرا که این روش‌ها ناپایداری مدل را از طریق افت عملکرد شناسایی می‌کنند. با این حال، روش‌های مرجع برای تشخیص ناپایداری به داده‌های آزمون برچسب‌دار نیاز دارند که خود موجب تأخیر در تشخیص می‌شود. همچنین، مسئله عدم توازن داده‌ها (یعنی تعداد نمونه‌های کلاس اقلیت به مراتب کمتر از نمونه‌های سالم کلاس اکثریت است) در JIT-SDP به عنوان یک چالش مطرح است؛ زیرا موجب کاهش قابلیت اطمینان معیار دقت کلی و عملکرد معیارهای وابسته به آستانه می‌شود. به همین دلیل، روش پیشنهادی روی داده‌های بازنمونه‌گیری‌شده نیز آزمایش و با روش‌های مرجع مبتنی بر معیارهای وابسته و مستقل از آستانه مقایسه می‌شود. نتایج حاصل از مطالعه روی ۲۰ مجموعه‌داده نشان می‌دهد که IID-JSDP قادر است ناپایداری مدل را در طول زمان با دقت بالا (به طور میانگین به دقت ۹۸٪) پیش‌بینی کند.
کلیدواژه‌ها

[1] X. Chen, Y. Zhao, Q. Wang, and Z. Yuan, "MULTI: Multi-objective effort-aware just-in-time software defect prediction," Information and Software Technology, vol. 93, pp. 1-13, 2018, doi: 10.1016/j.infsof.2017.08.004.
[2] G. G. Cabral, L. L. Minku, E. Shihab, and S. Mujahid, "Class imbalance evolution and verification latency in just-in-time software defect prediction," in 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE), 2019: IEEE, pp. 666-676, doi: 10.1109/ICSE.2019.00076.
[3] A. K. Gangwar, S. Kumar, and A. Mishra, "A Paired Learner-Based Approach for Concept Drift Detection and Adaptation in Software Defect Prediction," Applied Sciences, vol. 11, no. 14, p. 6663, 2021, doi: 10.3390/app11146663.
[4] F. Bayram, B. S. Ahmed, and A. Kassler, "From concept drift to model degradation: An overview on performance-aware drift detectors," Knowledge-Based Systems, vol. 245, p. 108632, 2022, doi: 10.1016/j.knosys.2022.108632.
[5] F. Hinder, V. Vaquet, and B. Hammer, "One or two things we know about concept drift—a survey on monitoring in evolving environments. Part A: detecting concept drift," Frontiers in Artificial Intelligence, vol. 7, p. 1330257, 2024, doi: 10.3389/frai.2024.1330257.
[6] S. Agrahari and A. K. Singh, "Concept drift detection in data stream mining: A literature review," Journal of King Saud University-Computer and Information Sciences, vol. 34, no. 10, pp. 9523-9540, 2022, doi: 10.1016/j.jksuci.2021.11.006.
[7] M. Jain, G. Kaur, and V. Saxena, "A K-Means clustering and SVM based hybrid concept drift detection technique for network anomaly detection," Expert Systems with Applications, vol. 193, p. 116510, 2022, doi: 10.1016/j.eswa.2022.116510.
[8] S. McIntosh and Y. Kamei, "Are fix-inducing changes a moving target? a longitudinal case study of just-in-time defect prediction," in Proceedings of the 40th International Conference on Software Engineering, 2018, pp. 560-560, doi: 10.1145/3180155.3182514.
[9] D. Lin, C. Tantithamthavorn, and A. E. Hassan, "The impact of data merging on the interpretation of cross-project just-in-time defect models," IEEE Transactions on Software Engineering, vol. 48, no. 8, pp. 2969-2986, 2021, doi: 10.1109/TSE.2021.3073920.
[10]     G. K. Rajbahadur, S. Wang, G. A. Oliva, Y. Kamei, and A. E. Hassan, "The impact of feature importance methods on the interpretation of defect classifiers," IEEE Transactions on Software Engineering, vol. 48, no. 7, pp. 2245-2261, 2021, doi: 10.1109/TSE.2021.3056941.
[11]     W. Zheng, T. Shen, X. Chen, and P. Deng, "Interpretability application of the Just-in-Time software defect prediction model," Journal of Systems and Software, vol. 188, p. 111245, 2022, doi: 10.1016/j.jss.2022.111245.
[12]     K. Fathi, M. Sadurski, T. Kleinert, and H. W. van de Venn, "Source component shift detection & classification for improved remaining useful life estimation in alarm-based predictive maintenance," in 2023 23rd international conference on control, automation and systems (iccas), 2023: IEEE, pp. 975-980, doi: 10.23919/ICCAS59377.2023.10316874.
[13]     B. Turhan, "On the dataset shift problem in software engineering prediction models," Empirical Software Engineering, vol. 17, pp. 62-74, 2012, doi: 10.1007/s10664-011-9182-8.
[14]     J. Lu, A. Liu, F. Dong, F. Gu, J. Gama, and G. Zhang, "Learning under concept drift: A review," IEEE transactions on knowledge and data engineering, vol. 31, no. 12, pp. 2346-2363, 2018, doi: 10.1109/TKDE.2018.2876857.
[15]     F. Dong, J. Lu, K. Li, and G. Zhang, "Concept drift region identification via competence-based discrepancy distribution estimation," in 2017 12th International Conference on Intelligent Systems and Knowledge Engineering (ISKE), 2017: IEEE, pp. 1-7, doi: 10.1109/ISKE.2017.8258734.
[16]     R. F. Haase and M. V. Ellis, "Multivariate analysis of variance," Journal of Counseling Psychology, vol. 34, no. 4, p. 404, 1987.
[17]     J. Gama, Knowledge discovery from data streams. CRC Press, 2010.
[18]     O. A. Mahdi, E. Pardede, N. Ali, and J. Cao, "Fast reaction to sudden concept drift in the absence of class labels," Applied Sciences, vol. 10, no. 2, p. 606, 2020, doi: 10.3390/app10020606.
[19]     J. Jiarpakdee, C. K. Tantithamthavorn, H. K. Dam, and J. Grundy, "An empirical study of model-agnostic techniques for defect prediction models," IEEE Transactions on Software Engineering, vol. 48, no. 1, pp. 166-185, 2020, doi: 10.1109/TSE.2020.2982385.
[20]     E. Štrumbelj, I. Kononenko, and M. R. Šikonja, "Explaining instance classifications with interactions of subsets of feature values," Data & Knowledge Engineering, vol. 68, no. 10, pp. 886-904, 2009, doi: 10.1016/j.datak.2009.01.004.
[21]     M. T. Ribeiro, S. Singh, and C. Guestrin, "" Why should i trust you?" Explaining the predictions of any classifier," in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, 2016, pp. 1135-1144, doi: 10.1145/2939672.2939778.
[22]     S. M. Lundberg and S.-I. Lee, "A unified approach to interpreting model predictions," Advances in neural information processing systems, vol. 30, 2017.
[23]     D. Vreš and M. Robnik-Šikonja, "Preventing deception with explanation methods using focused sampling," Data Mining and Knowledge Discovery, vol. 38, no. 5, pp. 3262-3307, 2024, doi: 10.1007/s10618-022-00900-w.
[24]     Y. Kamei, T. Fukushima, S. McIntosh, K. Yamashita, N. Ubayashi, and A. E. Hassan, "Studying just-in-time defect prediction using cross-project models," Empirical Software Engineering, vol. 21, pp. 2072-2106, 2016, doi: 10.1007/s10664-015-9400-x.
[25]     W. Li, W. Zhang, X. Jia, and Z. Huang, "Effort-aware semi-supervised just-in-time defect prediction," Information and Software Technology, vol. 126, p. 106364, 2020, doi: 10.1016/j.infsof.2020.106364.
[26]     G. J. Ross, N. M. Adams, D. K. Tasoulis, and D. J. Hand, "Exponentially weighted moving average charts for detecting concept drift," Pattern recognition letters, vol. 33, no. 2, pp. 191-198, 2012.
[27]     C. Tantithamthavorn, A. E. Hassan, and K. Matsumoto, "The impact of class rebalancing techniques on the performance and interpretation of defect prediction models," IEEE Transactions on Software Engineering, vol. 46, no. 11, pp. 1200-1219, 2020, doi: 10.1109/TSE.2018.2876537.
[28]     L. Pelayo and S. Dick, "Applying novel resampling strategies to software defect prediction," in NAFIPS 2007-2007 Annual meeting of the North American fuzzy information processing society, 2007: IEEE, pp. 69-72, doi: 10.1109/NAFIPS.2007.383813.
[29]     L. Chen, B. Fang, Z. Shang, and Y. Tang, "Tackling class overlap and imbalance problems in software defect prediction," Software Quality Journal, vol. 26, pp. 97-125, 2018, doi: 10.1007/s11219-016-9342-6.
[30]     O. I. Sheluhin and S. A. Sekretarev, "Concept drift detection in streaming classification of mobile application traffic," Automatic Control and Computer Sciences, vol. 55, pp. 253-262, 2021, doi: 10.1016/j.procs.2017.11.440.
[31]     S. Tabassum, L. L. Minku, and D. Feng, "Cross-project online just-in-time software defect prediction," IEEE Transactions on Software Engineering, vol. 49, no. 1, pp. 268-287, 2022, doi: 10.1109/TSE.2022.3150153.
[32]     Y. Gao, Y. Zhu, and Y. Zhao, "Dealing with imbalanced data for interpretable defect prediction," Information and software technology, vol. 151, p. 107016, 2022, doi: 10.1016/j.infsof.2022.107016.
[33]     A. Mockus and D. M. Weiss, "Predicting risk of software changes," Bell Labs Technical Journal, vol. 5, no. 2, pp. 169-180, 2000, doi: 10.1002/bltj.2229.
[34]     A. E. Hassan, "Predicting faults using the complexity of code changes," in 2009 IEEE 31st international conference on software engineering, 2009: IEEE, pp. 78-88, doi: 10.1109/ICSE.2009.5070510.
[35]     R. Purushothaman and D. E. Perry, "Toward understanding the rhetoric of small source code changes," IEEE Transactions on Software Engineering, vol. 31, no. 6, pp. 511-526, 2005, doi: 10.1109/TSE.2005.74.
[36]     P. J. Guo, T. Zimmermann, N. Nagappan, and B. Murphy, "Characterizing and predicting which bugs get fixed: an empirical study of microsoft windows," in Proceedings of the 32Nd ACM/IEEE International Conference on Software Engineering-Volume 1, 2010, pp. 495-504, doi: 10.1145/1806799.1806871.
[37]     L. Torgo and M. L. Torgo, "Package ‘dmwr’," Comprehensive R archive network, 2013.
[38]     A. L. Suárez-Cetrulo, D. Quintana, and A. Cervantes, "A survey on machine learning for recurring concept drifting data streams," Expert Systems with Applications, vol. 213, p. 118934, 2023, 10.1016/j.eswa.2022.118934.
[39]     A. Bifet, J. Read, B. Pfahringer, G. Holmes, and I. Žliobaitė, "CD-MOA: Change detection framework for massive online analysis," in International Symposium on Intelligent Data Analysis, 2013: Springer, pp. 92-103,  doi: 10.1007/978-3-642-41398-8_9.
[40]     J. Demšar and Z. Bosnić, "Detecting concept drift in data streams using model explanation," Expert Systems with Applications, vol. 92, pp. 546-559, 2018, doi: 10.1016/j.eswa.2017.10.003.
[41]     W. J. Conover and R. L. Iman, "Rank transformations as a bridge between parametric and nonparametric statistics," The American Statistician, vol. 35, no. 3, pp. 124-129, 1981.
[42]     R. C. Blair and J. J. Higgins, "The power of t and Wilcoxon statistics: A comparison," Evaluation Review, vol. 4, no. 5, pp. 645-656, 1980, doi: 10.1177/0193841X8000400506.