مجله علمی  رایانش نرم و فناوری اطلاعات

مجله علمی رایانش نرم و فناوری اطلاعات

انتخاب ویژگی‌های بهینه برای تشخیص حملات فیشینگ مبتنی بر معیارهای شباهت دامنه

نوع مقاله : مقاله پژوهشی فارسی

نویسندگان
گروه علوم کامپیوتر، دانشگاه سیستان و بلوچستان، زاهدان، ایران.
چکیده
امروزه رشد روزافزون فضای مجازی بر ابعاد مختلف زندگی انسان تأثیر گذاشته و مزیت‌های استفاده از این فناوری موجب شده است که بسیاری از کسب‌وکارها بر پایه تبادل اطلاعات در این فضا شکل بگیرند. با وجود تمام مزایای این فناوری، مخاطراتی نیز ممکن است افراد استفاده‌کننده از آن را تهدید کند که یکی از این مخاطرات، محتواهای تقلبی است که می‌تواند به‌عنوان ابزاری برای انجام حملات فیشینگ مورد استفاده قرار گیرد. از آن‌جا که سایت‌های فیشینگ معمولاً ظاهری مشابه وب‌سایت‌های قانونی دارند و محتوای آن‌ها را تکرار می‌کنند، تشخیص آن‌ها به‌دلیل گردآوری و تحلیل محتوای سایت‌ها پیچیده می‌شود. لذا در این مقاله، به‌طور خاص به ویژگی‌های مرتبط با دامنه سایت پرداخته شده است. این ایده که دامنه یک سایت فیشینگ سعی در ایجاد تشابه با دامنه سایت‌های معتبر و قانونی دارد، در نظر گرفته شده است. به همین منظور، یک مجموعه ویژگی مبتنی بر شباهت بر اساس نتایج موتور جستجوی گوگل معرفی شده است. این ویژگی‌ها به همراه ویژگی‌های دیگر که در دو گروه ویژگی‌های مبتنی بر وابستگی و ویژگی‌های مبتنی بر ارزش توانسته‌اند با استفاده از 127 ویژگی، دقت 99.26 درصد در شناسایی سایت‌های فیشینگ را به دست آورند. همچنین، با ارزیابی ویژگی‌ها با استفاده از معیارهای انتخاب ویژگی فیلتر، نشان داده شده است که 5 ویژگی مبتنی بر شباهت معرفی‌شده توانسته‌اند رتبه‌های دوم تا ششم بهترین ویژگی‌ها را از نظر معیارهای انتخاب ویژگی کسب کنند. در نهایت، با کاهش ابعاد مجموعه ویژگی‌ها به 54 ویژگی، دقت طبقه‌بندی به 99.47 درصد بهبود یافته است.
کلیدواژه‌ها

[1] A. Oest, Y. Safaei, A. Doupé, G.-J. Ahn, B. Wardman, and K. Tyers, “PhishFarm: A scalable framework for measuring the effectiveness of evasion techniques against browser phishing blacklists,” In 2019 IEEE Symposium on Security and Privacy (SP), pp. 1344–1361, IEEE, 2019.
[2] Y. Zhang, J. I. Hong, and L. F. Cranor, “Cantina: a content-based approach to detecting phishing websites,” In Proceedings of the 16th International Conference on World Wide Web, pp. 639–648, 2007.
[3] CVE Details, Google Chrome vulnerabilities, Retrieved from https://www.cvedetails.com/product/15031/Google-Chrome.html?vendor_id=1224, n.d.
[4] D.-J. Liu, G.-G. Geng, X.-B. Jin, and W. Wang, “An efficient multistage phishing website detection model based on the CASE feature framework: Aiming at the real web environment,” Journal of Computer Research and Development, Vol. 58, No. 10, pp. 2201, 2021.
[5] J. Zhou, H. Cui, X. Li, W. Yang, and X. Wu, “A novel phishing website detection model based on LightGBM and domain name features,” Symmetry, Vol. 15, No. 1, 180, 2023.
[6] A. Prasad and S. Chandra, “PhiUSIIL: A diverse security profile empowered phishing URL detection framework based on similarity index and incremental learning,” Computers & Security, Vol. 136, 103545, 2024.
[7] M. Wang, L. Song, L. Li, Y. Zhu, and J. Li, “Phishing webpage detection based on global and local visual similarity,” Expert Systems with Applications, Vol. 252, 124120, 2024.
[8] P. K. Mvula, P. Branco, G. V. Jourdan, and H. L. Viktor, “COVID-19 malicious domain names classification,” Expert Systems with Applications, Vol. 204, 117553, 2022.
[9] H. C. Sit, A. Esmradi, D. W. K. Yip, and P. Sun, “An effective and robust similarity-based phishing website detector in cyber-physical systems,” in Proceedings of the IEEE Conference on Communications and Network Security (CNS), pp. 1–6, 2025.
[10] A. S. Bozkir and M. Aydos, “LogoSENSE: A companion HOG based logo detection scheme for phishing web page and E-mail brand recognition,” Computers & Security, Vol. 95, 101855, 2020.
[11] B. Kılıç and B. Çeliktaş, “Phishing attack detection using multi-scale visual similarity analysis,” in Proc. 33rd Signal Processing and Communications Applications Conference (SIU), pp. 1–4, 2025.
[12] G. Varshney, A. Raj, D. Sangwan, S. Abuadbba, R. Mishra, and Y. Gao, “A login page transparency and visual similarity-based zero-day phishing defense protocol,” Computers & Security, 104598, 2025.
[13] S. Mousavi and M. Bahaghighat, “Phishing website detection: An in-depth investigation of feature selection and deep learning,” Expert Systems, Vol. 42, No. 3, e13824, 2025.
[14] W. Wei, Q. Ke, J. Nowak, M. Korytkowski, R. Scherer, and M. Woźniak, “Accurate and fast URL phishing detector: A convolutional neural network approach,” Computer Networks, Vol. 178, Art. no. 107275, 2020.
[15] M. A. Saeed, A. S. S. Balaid, and O. Bahaidara, “Phishing URL detection using deep learning: A CNN-based approach,” Journal of Science and Technology, Vol. 30, No. 9, 2025.
[16] M. Pranav, “URLTran: Improving phishing URL detection using transformers,” arXiv preprint, arXiv:2106.v2, 2017.
[17] S. Asiri, Y. Xiao, and T. Li, “PhishTransformer: A novel approach to detect phishing attacks using URL collection and transformer,” Electronics, Vol. 13, No. 1, Art. no. 30, 2023.
[18] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 785–794, ACM, 2016.
[19] G. Ke, Q. Meng, T. Finley, et al., “LightGBM: A highly efficient gradient boosting decision tree,” In Advances in Neural Information Processing Systems (NeurIPS), vol. 30, pp. 3146–3154, 2017.
[20] L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin, “CatBoost: Unbiased boosting with categorical features,” In Advances in Neural Information Processing Systems (NeurIPS), vol. 31, pp. 6638–6648, 2018.
[21] S. O. Arik and T. Pfister, “TabNet: Attentive interpretable tabular learning,” In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 8, pp. 6679–6687, AAAI Press, 2021.
[22] L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, Springer, 2001.
[23] Y. Freund and R. E. Schapire, “A decision-theoretic generalization of on-line learning and an application to boosting,” Journal of Computer and System Sciences, vol. 55, no. 1, pp. 119–139, 1997.
[24] J. R. Quinlan, “Induction of decision trees,” Machine Learning, vol. 1, no. 1, pp. 81–106, Springer, 1986.
[25] T. Cover and P. Hart, “Nearest neighbor pattern classification,” IEEE Transactions on Information Theory, vol. 13, no. 1, pp. 21–27, 1967.
[26] D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature, vol. 323, pp. 533–536, 1986.
[27] C. Cortes and V. Vapnik, “Support-vector networks,” Machine Learning, vol. 20, no. 3, pp. 273–297, Springer, 1995.
[28] I. Rish, “An empirical study of the naive Bayes classifier,” In IJCAI Workshop on Empirical Methods in Artificial Intelligence, pp. 41–46, 2001.
[29] D. R. Cox, “The regression analysis of binary sequences,” Journal of the Royal Statistical Society: Series B, vol. 20, no. 2, pp. 215–242, 1958.
[30] I. Guyon and A. Elisseeff, “An introduction to variable and feature selection,” Journal of Machine Learning Research, Vol. 3, pp. 1157–1182, 2003.
[31] D. C. Montgomery and G. C. Runger, Applied statistics and probability for engineers, 5th ed., Wiley, 2010.
[32] A. Agresti, Statistical inference, 2nd ed., Wiley, 2018.
[33] G. V. Bard, “Spelling-error tolerant, order-independent pass-phrases via the Damerau-Levenshtein string-edit distance metric,” In Proceedings of the Fifth Australasian Symposium on ACSW Frontiers, Vol. 68, pp. 117–124, 2007.
[34] M. Kiwi, M. Loebl, and J. Matoušek, “Expected length of the longest common subsequence for large alphabets,” Advances in Mathematics, Vol. 197, No. 2, pp. 480–498, 2005.
[35] J. Huang, Z. Fang, and H. Kasai, “LCS graph kernel based on Wasserstein distance in longest common subsequence metric space,” Signal Processing, Vol. 189, 108281, 2021.
[36] G. Kondrak, “N-gram similarity and distance,” In International Symposium on String Processing and Information Retrieval, pp. 115–126, Springer, Berlin, Heidelberg, 2005.
[37] N. Ohkura, M. Kiyomi, and K. Hirata, “The q-gram distance for ordered unlabeled trees,” In International Conference on Discovery Science, pp. 283–294, Springer, Berlin, Heidelberg, 2005.
[38] F. Rahutomo, T. Kitasuka, and M. Aritsugi, “Semantic cosine similarity,” In The 7th International Student Conference on Advanced Science and Technology (ICAST), Vol. 4, No. 1, 2012.
[39] W. Hidayat, E. Utami, and A. D. Hartanto, “Effect of Stemming Nazief & Adriani on the Ratcliff/Obershelp algorithm in identifying level of similarity between slang and formal words,” In 2020 3rd International Conference on Information and Communications Technology (ICOIACT), pp. 358–362, IEEE, 2020.
[40] feature23, StringSimilarity.NET, [Software]. GitHub. Retrieved from https://github.com/feature23/StringSimilarity.NET, n.d.
[41] G. Vrbančič, I. Fister Jr, and V. Podgorelec, “Datasets for phishing websites detection,” Data in Brief, Vol. 33, 106438, 2020.
[42] K. T. Chen, C. R. Huang, and C. S. Chen, “Fighting phishing with discriminative keypoint features,” IEICE Transactions on Information and Systems, Vol. 92, No. 5, pp. 1017–1030, 2009.
[43] K. L. Chiew, E. H. Chang, S. N. Sze, and W. K. Tiong, “Utilization of website logo for phishing detection,” Computers & Security, Vol. 54, pp. 16–26, 2015.
[44] X. Guang, O. Jason, P. Carolyn, and C. Lorrie, “CANTINA+: A feature-rich machine learning framework for detecting phishing websites,” ACM Transactions on Information and System Security, Vol. 14, No. 2, pp. 1–28, 2011.