Large-Scale Attribute-Object Compositions

doi:10.48550/arXiv.2105.11373

Large-Scale Attribute-Object Compositions

We study the problem of learning how to predict attribute-object compositions from images, and its generalization to unseen compositions missing from the training data. To the best of our knowledge, this is a first large-scale study of this problem, involving hundreds of thousands of compositions. We train our framework with images from Instagram using hashtags as noisy weak supervision. We make careful design choices for data collection and modeling, in order to handle noisy annotations and unseen compositions. Finally, extensive evaluations show that learning to compose classifiers outperforms late fusion of individual attribute and object predictions, especially in the case of unseen attribute-object pairs.

Publication:

arXiv e-prints

Pub Date:

May 2021

DOI:

10.48550/arXiv.2105.11373

arXiv:

arXiv:2105.11373

Bibcode:

2021arXiv210511373R

Keywords:

Computer Science - Computer Vision and Pattern Recognition

NASA/ADS

Large-Scale Attribute-Object Compositions

Abstract