HOPE-Net: A Graph-based Model for Hand-Object Pose Estimation

doi:10.48550/arXiv.2004.00060

HOPE-Net: A Graph-based Model for Hand-Object Pose Estimation

Hand-object pose estimation (HOPE) aims to jointly detect the poses of both a hand and of a held object. In this paper, we propose a lightweight model called HOPE-Net which jointly estimates hand and object pose in 2D and 3D in real-time. Our network uses a cascade of two adaptive graph convolutional neural networks, one to estimate 2D coordinates of the hand joints and object corners, followed by another to convert 2D coordinates to 3D. Our experiments show that through end-to-end training of the full network, we achieve better accuracy for both the 2D and 3D coordinate estimation problems. The proposed 2D to 3D graph convolution-based model could be applied to other 3D landmark detection problems, where it is possible to first predict the 2D keypoints and then transform them to 3D.

Publication:

arXiv e-prints

Pub Date:

March 2020

DOI:

10.48550/arXiv.2004.00060

arXiv:

arXiv:2004.00060

Bibcode:

2020arXiv200400060D

Keywords:

Computer Science - Computer Vision and Pattern Recognition

E-Print:

IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

NASA/ADS

HOPE-Net: A Graph-based Model for Hand-Object Pose Estimation

Abstract