Counterfactual Learning with General Data-generating Policies

doi:10.48550/arXiv.2212.01925

Counterfactual Learning with General Data-generating Policies

Off-policy evaluation (OPE) attempts to predict the performance of counterfactual policies using log data from a different policy. We extend its applicability by developing an OPE method for a class of both full support and deficient support logging policies in contextual-bandit settings. This class includes deterministic bandit (such as Upper Confidence Bound) as well as deterministic decision-making based on supervised and unsupervised learning. We prove that our method's prediction converges in probability to the true performance of a counterfactual policy as the sample size increases. We validate our method with experiments on partly and entirely deterministic logging policies. Finally, we apply it to evaluate coupon targeting policies by a major online platform and show how to improve the existing policy.

Publication:

arXiv e-prints

Pub Date:

December 2022

DOI:

10.48550/arXiv.2212.01925

arXiv:

arXiv:2212.01925

Bibcode:

2022arXiv221201925N

Keywords:

Computer Science - Machine Learning;
Computer Science - Artificial Intelligence;
Economics - Econometrics;
Statistics - Applications;
Statistics - Machine Learning

E-Print:

arXiv admin note: text overlap with arXiv:2104.12909

NASA/ADS

Counterfactual Learning with General Data-generating Policies

Abstract