مجله علمی  رایانش نرم و فناوری اطلاعات

مجله علمی رایانش نرم و فناوری اطلاعات

استفاده از شبکه‌های عصبی گرافی و مکانیزم توجه برای تشخیص حرکت انسان مبتنی بر داده‌های اسکلتی

نوع مقاله : مقاله پژوهشی فارسی

نویسندگان
دانشکده فنی و مهندسی، دانشگاه شهرکرد، شهرکرد، ایران.
چکیده
در سال‌های اخیر شبکه‌های کانولوشن گرافی (GCN) عملکرد قابل توجهی در زمینه‌ی تشخیص حرکت مبتنی بر اسکلت به دست آورده‌اند. روش‌های مبتنی بر GCN موجود، معمولاً توپولوژی‌های گرافی ثابت و یک فیلتر کانولوشنی زمانی ثابت را برای استخراج ویژگی‌های مکانی و زمانی یک حرکت اعمال می‌کنند. از آنجایی که یک حرکت انجام شده توسط انسان، از طریق بخش‌های مختلف بدن در حوزه‌های زمانی هماهنگ می‌شوند و ویژگی‌های مختلفی را در حوزه‌ی زمانی نشان می‌دهند، این کار باعث از دست رفتن اطلاعات زیادی برای یک حرکت می‌شود. برای پرداختن به این موضوع، در این مقاله یک شبکه‌ی عصبی گرافی مبتنی بر توجه (AT-AR) برای کشف ویژگی‌های متمایز از جنبه‌های مکانی و زمانی ارائه می‌کنیم. مدل پیشنهادی از یک کانولوشن SPG Net برای یادگیری ویژگی‌های مکانی – زمانی استفاده می‌کند.ما همچنین از دو ماژول توجه استفاده می‌کنیم. مکانیسم توجه STA، یک امتیاز توجه را با استفاده از ویژگی‌های زمانی ایجاد می‌کند، که می‌تواند همبستگی‌های زمانی یک حرکت را افزایش دهد و مکانیزم خودتوجه نیز مفاصلی را انتخاب می‌کند که برای تشخیص حرکت مهم‌تر هستند. این دو ماژول توجه در یک شبکه دوجریانی با هم ترکیب شده‌اند و با استفاده از ورودی یکسان در دیتاست NTU RGB+D 60 کار تشخیص حرکت مبتنی بر اسکلت را انجام می‌دهند.
کلیدواژه‌ها

[1] P. Yin, J. Ye, G. Lin, and Q. Wu, “Graph neural network for 6D object pose estimation,” Knowl.-Based Syst., vol. 218, Art. no. 106839, 2021, doi: 10.1016/j.knosys.2021.106839.
[2] D. Kong, Y. Bao, and W. Chen, “Collaborative learning based on centroid distance-vector for wearable devices,” Knowl.-Based Syst., vol. 194, Art. no. 105569, 2020, doi: 10.1016/j.knosys.2020.105569.
[3] M. Zhang, G. Tian, Y. Zhang, and P. Duan, “Service skill improvement for home robots: autonomous generation of action sequence based on reinforcement learning,” Knowl.-Based Syst., vol. 212, Art. no. 106605, 2021, doi: 10.1016/j.knosys.2020.106605.
[4] I. S. MacKenzie, Human-Computer Interaction: An Empirical Research Perspective, 2nd ed. Amsterdam, The Netherlands: Morgan Kaufmann, 2024, doi: 10.1016/C2020-0-02009-6.
[5] X. Guo, S. Guo, C. Wu, J. Li, C. Liu, and W. Chen, “Intelligent monitoring for safety-enhanced lithium-ion/sodium-ion batteries,” Adv. Energy Mater., vol. 13, no. 10, Art. no. 2300268, 2023, doi: 10.1002/aenm.202300268.
[6] C. Feichtenhofer, H. Fan, J. Malik, and K. He, “SlowFast networks for video recognition,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Seoul, South Korea, Oct. 2019, pp. 6202–6211, doi: 10.1109/ICCV.2019.00630.
[7] S. Baek, Z. Shi, M. Kawade, and T.-K. Kim, “Kinematic-layout-aware random forests for depth-based action recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Las Vegas, NV, USA, Jun. 2016, pp. 1468–1477, doi: 10.1109/CVPR.2016.160.
[8] Y. Liu, Z. Lu, J. Li, T. Yang, and C. Yao, “Global temporal representation based CNNs for infrared action recognition,” IEEE Signal Process. Lett., vol. 25, no. 6, pp. 927–931, Jun. 2018, doi: 10.1109/LSP.2018.2823654.
[9] P. Zhang, C. Lan, J. Xing, W. Zeng, J. Xue, and N. Zheng, “View adaptive recurrent neural networks for high performance human action recognition,” in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), Venice, Italy, Oct. 2017, pp. 2117–2126, doi: 10.1109/ICCV.2017.229.
[10] Z. Zhang, “Microsoft Kinect sensor and its effect,” IEEE MultiMedia, vol. 19, no. 2, pp. 4–10, Apr. 2012, doi: 10.1109/MMUL.2012.24.
[11] ASUS, “Xtion PRO LIVE,” 2011. [Online]. Available: https://www.asus.com/3D-Sensor/Xtion_PRO/. [Accessed: 03-May-2025].
[12] Z. Cao, G. Hidalgo, T. Simon, S.-E. Wei, and Y. Sheikh, “OpenPose: realtime multi-person 2D pose estimation using Part Affinity Fields,” arXiv preprint arXiv:1812.08008, Dec. 2018.
[13] H.-S. Fang, J. Li, H. Tang, C. Xu, H. Zhu, Y. Xiu, Y.-L. Li, and C. Lu, “AlphaPose: Whole-body regional multi-person pose estimation and tracking in real-time,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 8, pp. 7157–7173, Aug. 2022, doi: 10.1109/TPAMI.2022.3222784.
[14] Y. Liu, H. Zhang, D. Xu, and K. He, “Graph transformer network with temporal kernel attention for skeleton-based action recognition,” Knowl.-Based Syst., vol. 239, Art. no. 108146, Dec. 2022, doi: 10.1016/j.knosys.2022.108146.
[15] F. Han, B. Reily, W. Hoff, and H. Zhang, “Space-time representation of people based on 3D skeletal data: A review,” Comput. Vis. Image Underst., vol. 158, pp. 85–105, Jan. 2017, doi: 10.1016/j.cviu.2017.01.011.
[16] R. Hou and Z. Wang, “Self-attention based anchor proposal for skeleton-based action recognition,” arXiv preprint arXiv:2112.09413, Dec. 2021, doi: 10.48550/arXiv.2112.09413.
[17] Y. Du, Y. Fu, and L. Wang, “Skeleton-based action recognition with convolutional neural network,” in Proc. 3rd IAPR Asian Conf. Pattern Recognit. (ACPR), Kuala Lumpur, Malaysia, Nov. 2015, pp. 579–583, doi: 10.1109/ACPR.2015.7486569.
[18] H. Liu, J. Tu, and M. Liu, “Two-stream 3D convolutional neural network for skeleton-based action recognition,” arXiv preprint arXiv:1705.08106, May 2017.
[19] B. Li, Y. Dai, X. Cheng, H. Chen, Y. Lin, and M. He, “Skeleton-based action recognition using translation-scale invariant image mapping and multi-scale deep CNN,” in Proc. IEEE Int. Conf. Multimedia Expo Workshops (ICMEW), Hong Kong, China, Jul. 2017, pp. 601–604, doi: 10.1109/ICMEW.2017.8026282.
[20] T. S. Kim and A. Reiter, “Interpretable 3-D human action analysis with temporal convolutional networks,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), Snowbird, UT, USA, Jun. 2017, pp. 1623–1631, doi: 10.1109/CVPRW.2017.207.
[21] M. Liu, H. Liu, and C. Chen, “Enhanced skeleton visualization for view-invariant human action recognition,” Pattern Recognit., vol. 68, pp. 346–362, Aug. 2017, doi: 10.1016/j.patcog.2017.03.029.
[22] J. Liu, A. Shahroudy, D. Xu, and G. Wang, “Spatio-temporal LSTM with trust gates for 3-D human action recognition,” in Proc. Eur. Conf. Comput. Vis. (ECCV), Cham, Switzerland: Springer, 2016, pp. 816–833, doi: 10.1007/978-3-319-46487-9_50.
[23] I. Lee, D. Kim, S. Kang, and S. Lee, “Ensemble deep learning for skeleton-based action recognition using temporal sliding LSTM networks,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Venice, Italy, Oct. 2017, pp. 1012–1020, doi: 10.1109/ICCV.2017.115.
[24] P. Zhang, C. Lan, J. Xing, W. Zeng, J. Xue, and N. Zheng, “View adaptive recurrent neural networks for high performance human action recognition from skeleton data,” in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), Venice, Italy, Oct. 2017, pp. 2117–2126, doi: 10.1109/ICCV.2017.229.
[25]  C. Li, C. Xie, B. Zhang, J. Han, X. Zhen, and J. Chen, “Memory attention networks for skeleton-based action recognition,” IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 9, pp. 4800–4814, Sep. 2022, doi: 10.1109/TNNLS.2021.3061115.
[26] W. Li, L. Wen, M.-C. Chang, S.-N. Lim, and S. Lyu, “Adaptive RNN tree for large-scale human action recognition,” in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), Venice, Italy, Oct. 2017, pp. 1444–1452, doi: 10.1109/ICCV.2017.159.
[27] S. Yan, Y. Xiong, and D. Lin, “Spatial temporal graph convolutional networks for skeleton-based action recognition,” in Proc. AAAI Conf. Artif. Intell., vol. 32, no. 1, New Orleans, LA, USA, Feb. 2018, pp. 7444–7452.
[28] M. Wang, B. Ni, and X. Yang, “Learning multi-view interactional skeleton graph for action recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 6, pp. 6940–6954, Jun. 2023, doi: 10.1109/TPAMI.2020.3032738.
[29] T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” in Proc. Int. Conf. Learn. Represent. (ICLR), Toulon, France, Apr. 2017.
[30] S. Cho, M. Maqbool, F. Liu, and H. Foroosh, “Self-attention network for skeleton-based human action recognition,” in Proc. IEEE Winter Conf. Appl. Comput. Vis. (WACV), Snowmass Village, CO, USA, Jan. 2020, pp. 635–644, doi: 10.1109/WACV45572.2020.9093439.
[31] L. Shi, Y. Zhang, J. Cheng, and H. Lu, “Two-stream adaptive graph convolutional networks for skeleton-based action recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Long Beach, CA, USA, Jun. 2019, pp. 12026–12035, doi: 10.1109/CVPR.2019.01230.
[32] Y.-F. Song, Z. Zhang, and L. Wang, “Richly activated graph convolutional network for action recognition with incomplete skeletons,” in Proc. IEEE Int. Conf. Image Process. (ICIP), Taipei, Taiwan, Sep. 2019, pp. 1–5, doi: 10.1109/ICIP.2019.8802917.
[33] C. Si, W. Chen, W. Wang, L. Wang, and T. Tan, “An attention enhanced graph convolutional LSTM network for skeleton-based action recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Long Beach, CA, USA, Jun. 2019, pp. 1227–1236, doi: 10.1109/CVPR.2019.00132.
[34] Z. Liu, H. Zhang, Z. Chen, Z. Wang, and W. Ouyang, “Disentangling and unifying graph convolutions for skeleton-based action recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Seattle, WA, USA, Jun. 2020, pp. 140–149, doi: 10.1109/CVPR42600.2020.00023.
[35] A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “NTU RGB+D: A large scale dataset for 3D human activity analysis,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Las Vegas, NV, USA, Jun. 2016, pp. 1010–1019, doi: 10.1109/CVPR.2016.114.
[36] J. Lee, M. Lee, S. Cho, S. Woo, S. Jang, and S. Lee, “Leveraging spatio-temporal dependency for skeleton-based action recognition,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Paris, France, Oct. 2023, pp. 10255–10264, doi: 10.1109/ICCV51070.2023.00958.
[37] J. Lee, M. Lee, D. Lee, and S. Lee, “Hierarchically decomposed graph convolutional networks for skeleton-based action recognition,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Paris, France, Oct. 2023, pp. 10444–10453, doi: 10.1109/ICCV51070.2023.00958.
[38] Z. Qin, Y. Liu, P. Ji, D. Kim, L. Wang, R. I. McKay, S. Anwar, and T. Gedeon, “Fusing higher-order features in graph neural networks for skeleton-based action recognition,” IEEE Trans. Neural Netw. Learn. Syst., vol. 35, no. 4, pp. 4783–4797, Apr. 2024, doi: 10.1109/TNNLS.2022.3201518.
[39] J. Liu, X. Wang, C. Wang, Y. Gao, and M. Liu, “Temporal decoupling graph convolutional network for skeleton-based gesture recognition,” IEEE Trans. Multimedia, vol. 26, pp. 811–823, May 2023, doi: 10.1109/TMM.2023.3271811.
[40] Y. Zhou, Z.-Q. Cheng, C. Li, Y. Fang, Y. Geng, X. Xie, and M. Keuper, “Hypergraph transformer for skeleton-based action recognition,” arXiv preprint arXiv:2211.09590, Nov. 2022, doi: 10.48550/arXiv.2211.09590.