Detection of Human-Object Interactions with Deep Learning

Authors

DOI:

https://doi.org/10.67603/jaate.v1i02.10445

Keywords:

Human-object interaction detection, deep learning; Faster R-CNN, and SSD, U-Net++

Abstract

The In Human-object interaction (HOI) detection is a key area in computer vision, with applications in robotics, surveillance, autonomous vehicles, and human-computer interaction. The advent of deep learning, particularly the development of advanced architectures like R-CNN and YOLO, has significantly enhanced the accuracy and efficiency of object detection in complex environments. This paper explores the role of these models in improving HOI detection, focusing on their ability to identify and track objects in challenging scenarios. This research examines how deep learning models overcome challenges through advanced feature extraction, data augmentation,

and multi-scale processing. R-CNN and YOLO, two prominent models, each offer distinct advantages. R-CNN excels in accuracy due to its region -based approach, while YOLO is known for its speed and real-time process sing capabilities, making it ideal for applications requiring fast, on-the-fly detection.

Downloads

Download data is not yet available.

References

Y.LeCun, Bengio, Y. and Hinton, G. "Deep learning, (2015), Nature, vol. 521, no. 7553, pp. 436–444, doi: 10.1038/nature14539.

Y. Chao, Liu,Y. Liu, X. Zeng, H. and Deng, J. (2018), Learning to detect human-object interactions, in Proc. IEEE Winter Conf. on Applications of Computer Vision (WACV), pp. 381–389.

D. Papadopoulos, P. Uijlings, R. Keller, F. and Ferrari, V. (2017), Training object class detectors with click supervision, in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 5794–5803.

B. Wan, Zhou, T. Liu, Y. Ye, Q. and Jiao, J. (2019), Pose-aware multi-level feature network for human-object interaction detection, in Proc. IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 9469–9478.

, Z. Cao, Hidalgo, G. Simon, T. Wei, S. E. and Sheikh, Y. (2021), OpenPose: Realtime multi-person 2D pose estimation using Part Affinity Fields," IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 1, pp. 172–186, Jan.

G. Gkioxari, Girshick, Dollár, R. P. and He ,K. (2018), Detecting and recognizing human-object interactions, in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 8359–8367.

Y. Li, Zhou, S. Huang, Z. and Zhang, Y. (2021), HOI Transformer: Human-object interaction detection with transformers, arXiv preprint arXiv:2104.13682.

Y.-W. Chao,.Y. Liu, X. Liu, H. Zeng, and J. Deng, (2018), Learning to detect human-object interactions, in 2018 ieee winter conference on applications of computer vision (wacv). IEEE, pp. 381–389.

S. Qi, Wang, W. . Jia, Shen, BJ. and Zhu, S.-C. (2018), Learning human-object interactions by graph parsing neural networks, in Proceedings of the European conference on computer vision (ECCV), pp. 401–417.

B. Wan, Zhou, D. Liu,Y. Li, R. and He, X. (2019), Pose-aware multi-level feature network for human object interaction detection, in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9469–9478.

T. Gupta, Schwing, A. and Hoiem, D. (2019), No-frills human-object interac tion detection: Factorization, layout encodings, and training techniques,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9677–9685.

O. Ulutan, Iftekhar, A. and Manjunath, S. B. (2020), Vsgnet: Spatial attention network for detecting human object interactions using graph convolu tions, in Proceedings of the IEEE/CVF conference on computer vision and pattern recognitionpp. 13617–13626.

Y. Liu, Chen, Q. and Zisserman, A. (2020), “Amplifying key cues for human object-interaction detection,” in European Conference on Computer Vision. Springer, pp. 248–265.

D.-J. Kim,. Sun, X. Choi, Lin, J. S. and Kweon, I. S (2020), Detecting human object interactions with action co-occurrence priors, in Proceedings of the European Conference on Computer Vision (ECCV). Springer, pp. 718–736.

X. Zhong, Ding, C. Qu, X. and Tao, D. (2020), Polysemy deciphering network for human-object interaction detection,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, Proceedings, Part XX 16. pp. 69–85.

T. He, L. Gao, J. Song, and Y.-F. Li, (2021), Exploiting scene graphs for human-object interaction detection, in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15984–15993.

P. Zhou, , Hou, Q., and Cheng, M. (2022), Label decoupling framework for human-object interaction detection," IEEE Trans. Image Process., vol. 31, pp. 5389–5402,.

H. Xu, Zhu, Y. Choy, C. B. and Fei-Fei, L. (2019), Learning to detect and classify human-object interactions," Int. J. Comput. Vis., vol. 127, no. 3, pp. 258–275.

Q. Hou, P. Zhou., , & Cheng, M. M. (2021). HOI Analysis: Understanding human-object interaction by modeling object semantics. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 3027–3036). https://doi.org/10.1109/ICCV48922.2021.00302

R. Jain, Gandhi, H. and Bhattacharya, P. (2022), Lightweight human-object interaction detection for edge computing," IEEE Access, vol. 10, pp. 92846–92859,.

J. Redmon, and Farhadi, A. (2018), YOLOv3: An incremental improvement," arXiv preprint arXiv:1804.02767,

Z. Zhao, Zheng, Q. P. S. Xu, T., and Wu, X. (2019), Object detection with deep learning: A review,” IEEE Transactions on Neural Networks and Learning Systems, vol. 30, no. 11, pp. 3212–323.

Girshick, S. (2015), Fast R-CNN,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 2015, pp. 1440–1448, doi: 10.1109/ICCV. 169.

S. Ren, He, K. Girshick, R. and Sun, J. (2015), Faster R-CNN: Towards Real Time Object Detection with Region Proposal Networks,” in Advances in Neural Information Processing Systems (NeurIPS 2015).

J. Redmon, Divvala, S. Girshick, R. and Farhadi, A. ( 2016)“You Only Look Once: Unified, Real Time Object Detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, pp. 779–788, doi: 10.1109/CVPR.2016.91.

Downloads

Published

2026-06-30

How to Cite

Lalaoui, L., & Lakhnech, A. (2026). Detection of Human-Object Interactions with Deep Learning. Journal of Advanced Applied Technology and Engineering, 1(02), 52–59. https://doi.org/10.67603/jaate.v1i02.10445

Issue

Section

Articles