Cold-Start Reinforcement Learning with Softmax Policy Gradient
Abstract
Policy-gradient approaches to reinforcement learning have two common and undesirable overhead procedures, namely warm-start training and sample variance reduction. In this paper, we describe a reinforcement learning method based on a softmax value function that requires neither of these procedures. Our method combines the advantages of policy-gradient methods with the efficiency and simplicity of maximum-likelihood approaches. We apply this new cold-start reinforcement learning method in training sequence generation models for structured output prediction problems. Empirical evidence validates this method on automatic summarization and image captioning tasks.
- Publication:
-
arXiv e-prints
- Pub Date:
- September 2017
- DOI:
- 10.48550/arXiv.1709.09346
- arXiv:
- arXiv:1709.09346
- Bibcode:
- 2017arXiv170909346D
- Keywords:
-
- Computer Science - Machine Learning
- E-Print:
- Conference on Neural Information Processing Systems 2017. Main paper and supplementary material