DiME: Maximizing Mutual Information by a Difference of Matrix-Based Entropies

doi:10.48550/arXiv.2301.08164

DiME: Maximizing Mutual Information by a Difference of Matrix-Based Entropies

We introduce an information-theoretic quantity with similar properties to mutual information that can be estimated from data without making explicit assumptions on the underlying distribution. This quantity is based on a recently proposed matrix-based entropy that uses the eigenvalues of a normalized Gram matrix to compute an estimate of the eigenvalues of an uncentered covariance operator in a reproducing kernel Hilbert space. We show that a difference of matrix-based entropies (DiME) is well suited for problems involving the maximization of mutual information between random variables. While many methods for such tasks can lead to trivial solutions, DiME naturally penalizes such outcomes. We compare DiME to several baseline estimators of mutual information on a toy Gaussian dataset. We provide examples of use cases for DiME, such as latent factor disentanglement and a multiview representation learning problem where DiME is used to learn a shared representation among views with high mutual information.

Publication:

arXiv e-prints

Pub Date:

January 2023

DOI:

10.48550/arXiv.2301.08164

arXiv:

arXiv:2301.08164

Bibcode:

2023arXiv230108164S

Keywords:

Computer Science - Machine Learning;
Computer Science - Information Theory

NASA/ADS

DiME: Maximizing Mutual Information by a Difference of Matrix-Based Entropies

Abstract