Detecting (Un)Important Content for Single-Document News Summarization

doi:10.48550/arXiv.1702.07998

Detecting (Un)Important Content for Single-Document News Summarization

We present a robust approach for detecting intrinsic sentence importance in news, by training on two corpora of document-summary pairs. When used for single-document summarization, our approach, combined with the "beginning of document" heuristic, outperforms a state-of-the-art summarizer and the beginning-of-article baseline in both automatic and manual evaluations. These results represent an important advance because in the absence of cross-document repetition, single document summarizers for news have not been able to consistently outperform the strong beginning-of-article baseline.

Publication:

arXiv e-prints

Pub Date:

February 2017

DOI:

10.48550/arXiv.1702.07998

arXiv:

arXiv:1702.07998

Bibcode:

2017arXiv170207998Y

Keywords:

Computer Science - Computation and Language

E-Print:

Accepted By EACL 2017

ADS

Detecting (Un)Important Content for Single-Document News Summarization

Abstract