A Unified Maximum Likelihood Approach for Optimal Distribution Property Estimation
Abstract
The advent of data science has spurred interest in estimating properties of distributions over large alphabets. Fundamental symmetric properties such as support size, support coverage, entropy, and proximity to uniformity, received most attention, with each property estimated using a different technique and often intricate analysis tools. We prove that for all these properties, a single, simple, plug-in estimator---profile maximum likelihood (PML)---performs as well as the best specialized techniques. This raises the possibility that PML may optimally estimate many other symmetric properties.
- Publication:
-
arXiv e-prints
- Pub Date:
- November 2016
- DOI:
- 10.48550/arXiv.1611.02960
- arXiv:
- arXiv:1611.02960
- Bibcode:
- 2016arXiv161102960A
- Keywords:
-
- Computer Science - Information Theory;
- Computer Science - Data Structures and Algorithms;
- Computer Science - Machine Learning