مجله علمی  رایانش نرم و فناوری اطلاعات

مجله علمی رایانش نرم و فناوری اطلاعات

تقویت عملکرد پیش‌بینی برای یک مطالعه موردی در بیوانفورماتیک با استفاده از یک پیش‌بینی کننده مبتنی بر یادگیری (برهمکنش‌های پروتئین-پپتید در سطح باقیمانده‌های پیوندی)

نوع مقاله : مقاله پژوهشی فارسی

نویسندگان
گروه مهندسی کامپیوتر و فناوری اطلاعات، دانشکده فنی و مهندسی، دانشگاه رازی، کرمانشاه، ایران.
چکیده
پیش‌بینی برهمکنش‌های پروتئین-پپتید، در درک رفتارهای غیرطبیعی سلولی، تبیین عملکردهای بیولوژیکی، طراحی دارو و راهبردهای درمانی نقش مهمی دارد. روش‌های تجربی و آزمایشگاهی برای تشخیص باقیمانده‌های درگیر در برهمکنش پروتئین-پپتید، با محدودیت‌هایی مانند: هزینه‌های بالای نیروی انسانی، زمان‌بر بودن آزمایش‌ها و وابستگی به ابزارها و تجهیزات گران قیمت، همراه هستند. برای غلبه بر این محدودیت‌ها، یک پیش‌بینی‌کننده محاسباتی مبتنی بر بر یادگیری عمیق (واحد بازگشتی گیت‌دار عمیق) و یادگیری ماشین (ماشین بردار پشتیبان بهبود یافته) طراحی و پیاده‌سازی شده است. این مدل پیشنهادی، از انواع خصیصه‌های مستخرج شده از ساختارهای پروتئینی شامل ویژگی‌های تکاملی، مبتنی بر توالی، مبتنی بر ساختار و فیزیکوشیمیایی استفاده می‌کند. در مدل ترکیبی توسعه‌یافته از دو تکنیک نمونه‌برداری کاهشی و افزایشی برای مدیریت بیش پیش‌بینی کلاس غیرپیوندی ناشی از عدم توازن ذاتی داده‌های پروتئینی، استفاده شده است. عملکرد مدل پیشنهادی با دو مجموعه داده استاندارد (Sparks,Wei) برگرفته از پایگاه داده BioLip، سنجیده شده و در مقایسه با مدل‌های محاسباتی رقیب، موفق به أخذ نتایج مطلوب‌تر شده است. این نتایج شامل بهینگی از نظر معیارهای دقت (تقریباً1درصد)، اندازه‌گیری-اف (حداقل5/21 درصد)، تعادل مابین حساسیت و خاصگی (تقریباً 1/2درصد)، فاکتور غنی‌سازی، نقشه گرمایی و نرخ موفقیت براساس دامنه‌های پروتئینی است.
کلیدواژه‌ها

[1]     M. Y. Cao, S. Zainudin, and K. M. Daud, "Protein features fusion using attributed network embedding for predicting protein-protein interaction," BMC genomics, vol. 25, no. 1, p. 466, 2024, doi: 10.1186/s12864-024-10361-8.
 [2]    Z. Hekmati, J. Zahiri, and A. Aalami, "Computational prediction of protein–protein interactions’ network in Arabidopsis thaliana," Acta Physiologiae Plantarum, vol. 45, no. 12, p. 142 ,2023, doi: 10.1007/s11738-023-03623-7.
[3]     S. Shafiee, A. Fathi, and F. Abdali-Mohammadi, "A Review of the Uses of Artificial Intelligence in Protein Research," in the Fourth National Conference on Proteins and Peptide science, University of Isfahan,Isfahan,2019.
[4]     M. H. Viet et al., "In silico and in vitro study of binding affinity of tripeptides to amyloid β fibrils: implications for Alzheimer’s disease," The Journal of Physical Chemistry B, vol. 119, no. 16, pp. 5145-5155, 2015, doi: 10.1021/acs.jpcb.5b00006.
[5]     P. Scheltens et al., "Alzheimer's disease," The Lancet, vol. 397, no. 10284, pp. 1577-1590, 2021, doi: 10.1016/S0140-6736(20)32205-4.
[6]     E. Giusto, T. A. Yacoubian, E. Greggio, and L. Civiero, "Pathways to Parkinson’s disease: a spotlight on 14-3-3 proteins," npj Parkinson's Disease, vol. 7, no. 1, p. 85 2021, doi: 10.1038/s41531-021-00230-6.
[7]     B. R. Bloem, M. S. Okun, and C. Klein, "Parkinson's disease," The Lancet, vol. 397, no. 10291, pp. 2284-2303, 2021, doi: 10.1016/S0140-6736(21)00218-X.
[8]     E. C. Hutchinson, "Influenza virus," Trends in microbiology, vol. 26, no. 9, pp. 809-810, 2018, doi: 10.1016/j.tim.2018.05.013.
[9]     A. Fathi and R. Sadeghi, "A genetic programming method for feature mapping to improve prediction of HIV-1 protease cleavage site," Applied Soft Computing, vol. 72, pp. 56-64, 2018, doi: 10.1016/j.asoc.2018.06.045.
[10]   L. Yang et al., "COVID-19: immunopathogenesis and Immunotherapeutics," Signal transduction and targeted therapy, vol. 5, no. 1, p. 128, 2020, doi: 10.1038/s41392-020-00243-2.
[11]   Z. Zhao, Z. Peng, and J. Yang, "Improving sequence-based prediction of protein–peptide binding residues by introducing intrinsic disorder and a consensus method," Journal of Chemical Information and Modeling, vol. 58, no. 7, pp. 1459-1468, 2018, doi: 10.1021/acs.jcim.8b00019.
[12]   S. Shafiee, A. Fathi, and F. A. Mohammadi, "Prediction of protein–peptide binding residues using classification algorithms," In 2020 IEEE 20th International Conference on Bioinformatics and Bioengineering (BIBE), 2020,doi: 10.1109/BIBE50027.2020.00013.
[13]   S. Shafiee and A. Fathi, "Prediction of protein–peptide-binding amino acid residues regions using machine learning algorithms," In 2021 26th International Computer Conference, Computer Society of Iran (CSICC), 2021,doi: 10.1109/CSICC52343.2021.9420568.
[14]   E. Petsalaki, A. Stark, E. García-Urdiales, and R. B. Russell, "Accurate prediction of peptide binding sites on protein surfaces," PLoS computational biology, vol. 5, no. 3, p. e1000335, 2009, doi: 10.1371/journal.pcbi.1000335.
[15]   A. Lavi et al., "Detection of peptidebinding sites on protein surfaces: The first step toward the modeling and targeting of peptidemediated interactions," Proteins: Structure, Function, and Bioinformatics, vol. 81, no. 12, pp. 2096-2105, 2013, doi: 10.1002/prot.24422.
[16]   T. Kambe, B. E. Correia, M. J. Niphakis, and B. F. Cravatt, "Mapping the protein interaction landscape for fully functionalized small-molecule probes in human cells," Journal of the American Chemical Society, vol. 136, no. 30, pp. 10777-10782, 2014, doi: 10.1021/ja505517t.
[17]   H. Lee, L. Heo, M. S. Lee, and C. Seok, "GalaxyPepDock: a protein–peptide docking tool based on interaction similarity and energy optimization," Nucleic acids research, vol. 4, no. W1, pp. W431-W435, 2015, doi: 10.1093/nar/gkv495.
[18]   D. E. Shaw et al., "Atomic-level characterization of the structural dynamics of proteins," Science, vol. 330, no. 6002, pp. 341-346, 2010, doi: 10.1126/science.1187409.
[19]   G. Taherzadeh, Y. Yang, T. Zhang, A. W. C. Liew, and Y. Zhou, "Sequencebased prediction of proteinpeptide binding sites using support vector machine," Journal of computational chemistry, vol. 37, no. 13, pp. 1223-1229, 2016, doi: 10.1002/jcc.24314.
[20]   S. Shafiee and A. Fathi, "Combination of genetic programming and support vector machine-based prediction of protein-peptide binding sites with sequence and structure-based features," Journal of Computing and Security, vol. 8, no. 1, pp. 45-63, 2021, doi: 10.22108/jcs.2021.126817.1062.
[21]   S. Shafiee, A. Fathi, and G. Taherzadeh, "Spppred: Sequence-based protein-peptide binding residue prediction using genetic programming and ensemble learning," IEEE/ACM Transactions on Computational Biology and Bioinformatics, vol. 20, no. 3, pp. 2029-2040, 2022, doi: 10.1109/tcbb.2022.3230540.
[22]   W. Wardah et al., "Predicting protein-peptide binding sites with a deep convolutional neural network," Journal of Theoretical Biology, vol. 496, p. 110278, 2020, doi: 10.1016/j.jtbi.2020.110278.
[23]   O. Abdin, S. Nim, H. Wen, and P. M. Kim, "PepNN: a deep attention model for the identification of peptide binding sites," Communications biology, vol. 5, no. 1, p. 503, 2022, doi: 10.1038/s42003-022-03445-2.
[24]   R. Wang, J. Jin, Q. Zou, K. Nakai, and L. Wei, "Predicting protein–peptide binding residues via interpretable deep learning," Bioinformatics, vol. 38, no. 13, pp. 3351-3360, 2022, doi: 10.1093/bioinformatics/btac352.
[25]   A. Chandra, A. Sharma, I. Dehzangi, T. Tsunoda, and A. Sattar, "PepCNN deep learning tool for predicting peptide binding residues in proteins using sequence, structural, and language model features," Scientific Reports, vol. 13, no. 1, p. 20882, 2023, doi: 10.1038/s41598-023-47624 5.35.
[26]   zhanglab. "BioLip 2 for Ligand-protein binding database." http://zhanglab.ccmb.med.umich.edu/BioLiP (accessed 2024).
[27]   S. F. Altschul, W. Gish, W. Miller, E. W. Myers, and D. J. Lipman, "Basic local alignment search tool," Journal of molecular biology, vol. 215, no. 3, pp. 403-410, 1990, doi: 10.1016/S0022-2836(05)80360-2.
[28]   E. Faraggi, T. Zhang, Y. Yang, L. Kurgan, and Y. Zhou, "SPINE X: improving protein secondary structure prediction by multistep learning coupled with prediction of solvent accessible surface area and backbone torsion angles," Journal of computational chemistry, vol. 33, no. 3, pp. 259-267, 2012, doi: 10.1002/jcc.21968.
[29]   S. Iqbal, A. Mishra, and M. T. Hoque, "Improved prediction of accessible surface area results in efficient energy function application," Journal of theoretical biology, vol. 380, pp. 380-391, 2015, doi: 10.1016/j.jtbi.2015.06.012.
[30]   C. Fang, Y. Shang, and D. Xu, "Prediction of protein backbone torsion angles using deep residual inception neural networks," IEEE/ACM transactions on computational biology and bioinformatics, vol. 16, no. 3, pp. 1020-1028, 2018, doi: 10.1109/TCBB.2018.2814586.
[31]   N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: synthetic minority over-sampling technique," Journal of artificial intelligence research, vol. 16, pp. 321-357, 2002, doi: 10.1613/jair.953.
[32]   S. J. Yen and Y. S. Lee, "Under-sampling approaches for improving prediction of the minority class in an imbalanced dataset," In Lecture Notes in Control and Information Science , 2006, doi: 10.1007/978-1-4939-6406-2_6.
[33]   N. Sapoval, A. Aghazadeh, M. G. Nute, and et al, "Current progress and open challenges for applying deep learning across the biosciences," Nat Commun, vol. 13, p. 1728, 2022, doi: 10.1038/s41467-022-29268-7.
[34]   A. Biegert, C. Mayer, M. Remmert, J. Soding, and A. Lupas, "The MPI Bioinformatics Toolkit for protein sequence analysis," Nucleic acids research, vol. 2, no. suppl_2, pp. W335-W339, 2006, doi: 10.1093/nar/gkw348.
[35]   G. Taherzadeh, Y. Zhou, A. W. C. Liew, and Y. Yang, "Structure-based prediction of protein–peptide binding regions using Random Forest," Bioinformatics, vol. 34, no. 3, pp. 477-484, 2018, doi: 10.1093/bioinformatics/btx614.
[36]   S. Y. Novak, "On the T-test," Statistics & Probability Letters, vol. 189, 2022, doi: 1016/j.spl.2022.109562.
[37]  A. A. Mamun, A. Enan, D.A. Indah, J. Mwakalonge, G. Comert, and M. Chowdhury, "Crash severity risk modeling strategies under data imbalance," arxiv.org/abs/2412.02094, 2025,doi: 10.48550/arXiv.2412.02094.
[38]  X. Li, B. Cao, H. Ding, J. Mwakalonge, N. Kang, & T. Song, "PepPFN:protein-peptide binding residues prediction via pre-trained module-based Fourier Network," In 2024  IEEE Conference on Artificial Intelligence (CAI), Singapore, 2024,doi: 10.1109/CAI59869.2024.00195.
[39]  J. Hu, K.X Chen, B. Rao, J. YuanNi, M.A. Thafar, S. Albaradei & M. Arif, "Protein-Peptide Binding Residue Prediction Based on Protein Language Models and Cross-Attention Mechanism". Analytical Biochemistry, vol. 694 ,p. 115637, 2024,doi: 10.1016/j.ab.2024.115637.
[40]  Y. Yang, Y. Hua, W. Zhang, & X. Song, "DPPBP: Dual-stream Protein-Peptide Binding Sites Prediction Based on Region Detection". In 2025 International Conference on Intelligence , 2025, doi: 10.65286/icic.v21i2.56686.
[41]  D. Chicco, G. Jurman, "The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation", BMC Genomics ,vol.21 ,no.1, p.6 ,2020, doi:10.101186/s12864-019-6413-7.