On Design Choices in Similarity-Preserving Sparse Randomized Embeddings

doi:10.48550/arXiv.2501.14741

On Design Choices in Similarity-Preserving Sparse Randomized Embeddings

Expand & Sparsify is a principle that is observed in anatomically similar neural circuits found in the mushroom body (insects) and the cerebellum (mammals). Sensory data are projected randomly to much higher-dimensionality (expand part) where only few the most strongly excited neurons are activated (sparsify part). This principle has been leveraged to design a FlyHash algorithm that forms similarity-preserving sparse embeddings, which have been found useful for such tasks as novelty detection, pattern recognition, and similarity search. Despite its simplicity, FlyHash has a number of design choices to be set such as preprocessing of the input data, choice of sparsifying activation function, and formation of the random projection matrix. In this paper, we explore the effect of these choices on the performance of similarity search with FlyHash embeddings. We find that the right combination of design choices can lead to drastic difference in the search performance.

Publication:

arXiv e-prints

Pub Date:

December 2024

DOI:

10.48550/arXiv.2501.14741

arXiv:

arXiv:2501.14741

Bibcode:

2025arXiv250114741K

Keywords:

Computer Science - Neural and Evolutionary Computing;
Computer Science - Machine Learning;
Quantitative Biology - Neurons and Cognition

E-Print:

8 pages, 4 figures

ADS

On Design Choices in Similarity-Preserving Sparse Randomized Embeddings

Abstract