Improving Image Captioning by Leveraging Knowledge Graphs

doi:10.48550/arXiv.1901.08942

Improving Image Captioning by Leveraging Knowledge Graphs

We explore the use of a knowledge graphs, that capture general or commonsense knowledge, to augment the information extracted from images by the state-of-the-art methods for image captioning. The results of our experiments, on several benchmark data sets such as MS COCO, as measured by CIDEr-D, a performance metric for image captioning, show that the variants of the state-of-the-art methods for image captioning that make use of the information extracted from knowledge graphs can substantially outperform those that rely solely on the information extracted from images.

Publication:

arXiv e-prints

Pub Date:

January 2019

DOI:

10.48550/arXiv.1901.08942

arXiv:

arXiv:1901.08942

Bibcode:

2019arXiv190108942Z

Keywords:

Computer Science - Computer Vision and Pattern Recognition

E-Print:

Accepted by WACV'19

NASA/ADS

Improving Image Captioning by Leveraging Knowledge Graphs

Abstract