Learning to Remember Translation History with a Continuous Cache

doi:10.48550/arXiv.1711.09367

Learning to Remember Translation History with a Continuous Cache

Existing neural machine translation (NMT) models generally translate sentences in isolation, missing the opportunity to take advantage of document-level information. In this work, we propose to augment NMT models with a very light-weight cache-like memory network, which stores recent hidden representations as translation history. The probability distribution over generated words is updated online depending on the translation history retrieved from the memory, endowing NMT models with the capability to dynamically adapt over time. Experiments on multiple domains with different topics and styles show the effectiveness of the proposed approach with negligible impact on the computational cost.

Publication:

arXiv e-prints

Pub Date:

November 2017

DOI:

10.48550/arXiv.1711.09367

arXiv:

arXiv:1711.09367

Bibcode:

2017arXiv171109367T

Keywords:

Computer Science - Computation and Language

E-Print:

Accepted by TACL 2018

NASA/ADS

Learning to Remember Translation History with a Continuous Cache

Abstract