Accurate Shapley Values for explaining tree-based models

doi:10.48550/arXiv.2106.03820

Accurate Shapley Values for explaining tree-based models

Shapley Values (SV) are widely used in explainable AI, but their estimation and interpretation can be challenging, leading to inaccurate inferences and explanations. As a starting point, we remind an invariance principle for SV and derive the correct approach for computing the SV of categorical variables that are particularly sensitive to the encoding used. In the case of tree-based models, we introduce two estimators of Shapley Values that exploit the tree structure efficiently and are more accurate than state-of-the-art methods. Simulations and comparisons are performed with state-of-the-art algorithms and show the practical gain of our approach. Finally, we discuss the limitations of Shapley Values as a local explanation. These methods are available as a Python package.

Publication:

arXiv e-prints

Pub Date:

June 2021

DOI:

10.48550/arXiv.2106.03820

arXiv:

arXiv:2106.03820

Bibcode:

2021arXiv210603820A

Keywords:

Statistics - Machine Learning;
Computer Science - Machine Learning

E-Print:

Accepted at the 25th International Conference on Artificial Intelligence and Statistics (AISTATS), 2022. V2: The section on Active Shapley Values has been removed in this updated version

NASA/ADS

Accurate Shapley Values for explaining tree-based models

Abstract