Efficient Document Indexing Using Pivot Tree

doi:10.48550/arXiv.1605.06693

Efficient Document Indexing Using Pivot Tree

We present a novel method for efficiently searching top-k neighbors for documents represented in high dimensional space of terms based on the cosine similarity. Mostly, documents are stored as bag-of-words tf-idf representation. One of the most used ways of computing similarity between a pair of documents is cosine similarity between the vector representations, but cosine similarity is not a metric distance measure as it doesn't follow triangle inequality, therefore most metric searching methods can not be applied directly. We propose an efficient method for indexing documents using a pivot tree that leads to efficient retrieval. We also study the relation between precision and efficiency for the proposed method and compare it with a state of the art in the area of document searching based on inner product.

Publication:

arXiv e-prints

Pub Date:

May 2016

DOI:

10.48550/arXiv.1605.06693

arXiv:

arXiv:1605.06693

Bibcode:

2016arXiv160506693S

Keywords:

Computer Science - Information Retrieval;
Computer Science - Machine Learning

E-Print:

6 Pages, 2 Figures

NASA/ADS

Efficient Document Indexing Using Pivot Tree

Abstract