Annotation-free Learning of Deep Representations for Word Spotting using Synthetic Data and Self Labeling

doi:10.48550/arXiv.2003.01989

Annotation-free Learning of Deep Representations for Word Spotting using Synthetic Data and Self Labeling

Word spotting is a popular tool for supporting the first exploration of historic, handwritten document collections. Today, the best performing methods rely on machine learning techniques, which require a high amount of annotated training material. As training data is usually not available in the application scenario, annotation-free methods aim at solving the retrieval task without representative training samples. In this work, we present an annotation-free method that still employs machine learning techniques and therefore outperforms other learning-free approaches. The weakly supervised training scheme relies on a lexicon, that does not need to precisely fit the dataset. In combination with a confidence based selection of pseudo-labeled training samples, we achieve state-of-the-art query-by-example performances. Furthermore, our method allows to perform query-by-string, which is usually not the case for other annotation-free methods.

Publication:

arXiv e-prints

Pub Date:

March 2020

DOI:

10.48550/arXiv.2003.01989

arXiv:

arXiv:2003.01989

Bibcode:

2020arXiv200301989W

Keywords:

Computer Science - Computer Vision and Pattern Recognition

E-Print:

Accepted to Workshop on Document Analysis Systems (DAS) 2020

NASA/ADS

Annotation-free Learning of Deep Representations for Word Spotting using Synthetic Data and Self Labeling

Abstract