Generalizable Single-view Object Pose Estimation by Two-side Generating and Matching

doi:10.48550/arXiv.2411.15860

Generalizable Single-view Object Pose Estimation by Two-side Generating and Matching

In this paper, we present a novel generalizable object pose estimation method to determine the object pose using only one RGB image. Unlike traditional approaches that rely on instance-level object pose estimation and necessitate extensive training data, our method offers generalization to unseen objects without extensive training, operates with a single reference image of the object, and eliminates the need for 3D object models or multiple views of the object. These characteristics are achieved by utilizing a diffusion model to generate novel-view images and conducting a two-sided matching on these generated images. Quantitative experiments demonstrate the superiority of our method over existing pose estimation techniques across both synthetic and real-world datasets. Remarkably, our approach maintains strong performance even in scenarios with significant viewpoint changes, highlighting its robustness and versatility in challenging conditions. The code will be re leased at https://github.com/scy639/Gen2SM.

Publication:

arXiv e-prints

Pub Date:

November 2024

DOI:

10.48550/arXiv.2411.15860

arXiv:

arXiv:2411.15860

Bibcode:

2024arXiv241115860S

Keywords:

Computer Science - Computer Vision and Pattern Recognition

E-Print:

Accepted by WACV 2025, not published yet

NASA/ADS

Generalizable Single-view Object Pose Estimation by Two-side Generating and Matching

Abstract