Active Perception with Initial-State Uncertainty: A Policy Gradient Method

doi:10.48550/arXiv.2409.16439

Active Perception with Initial-State Uncertainty: A Policy Gradient Method

This paper studies the synthesis of an active perception policy that maximizes the information leakage of the initial state in a stochastic system modeled as a hidden Markov model (HMM). Specifically, the emission function of the HMM is controllable with a set of perception or sensor query actions. Given the goal is to infer the initial state from partial observations in the HMM, we use Shannon conditional entropy as the planning objective and develop a novel policy gradient method with convergence guarantees. By leveraging a variant of observable operators in HMMs, we prove several important properties of the gradient of the conditional entropy with respect to the policy parameters, which allow efficient computation of the policy gradient and stable and fast convergence. We demonstrate the effectiveness of our solution by applying it to an inference problem in a stochastic grid world environment.

Publication:

arXiv e-prints

Pub Date:

September 2024

DOI:

10.48550/arXiv.2409.16439

arXiv:

arXiv:2409.16439

Bibcode:

2024arXiv240916439S

Keywords:

Electrical Engineering and Systems Science - Systems and Control

ADS

Active Perception with Initial-State Uncertainty: A Policy Gradient Method

Abstract