Summarization
Summarization is when a model turns a long text into a shorter version that keeps the important information.
Summarization means turning a long text into a shorter one that keeps what matters. There are two main styles. Extractive summaries pick out the most relevant parts of the original. Abstractive summaries write new sentences, the way you would retell a film to a friend.
But abstractive summaries have a known weak spot. In one study, reviewers found a lot of made-up content in summaries from every model they tested. So a good summary is short and also faithful. Faithful means every point in it is really in the source.
People often score summaries with ROUGE. It is a tool that counts how much a summary overlaps with one written by an expert. But research found that other checks match faithfulness better than standard scores like this.
Long documents take time to read, and a short version is hard to get right.
Follow one long document through a model that summarizes it.
- 1 · askTell the model which details the summary should include.
- 2 · splitIf the text is too long for the model, break it into smaller chunks.
- 3 · condenseThe model writes a shorter version of each chunk.
- 4 · combineThe chunk summaries are merged into one summary of the whole document.
- 5 · checkIt is wise to compare the summary with the source, because models can add content that is not in the input.
A summary is only useful if it is faithful to the source.
| Who | What they ask | What it works with |
|---|---|---|
| Lawyer | “What are the key terms in this lease?” | Parties, dates, contract terms and specific clauses |
| Legal team | “What is in this big pile of contracts?” | Short summaries to speed up due diligence |
- It cuts a long text down to its important points.
- It can follow a set format, so many summaries are easy to compare.
- Chunking lets it handle documents too long to process in one go.
- It can include content that is not faithful to the original.
- Early neural summarizers could repeat themselves or get facts wrong.
- Research found other checks match faithfulness better than standard overlap scores.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsSummarization, Hugging Face · read 28 Sept 2026
- docsLegal summarization, Anthropic · read 28 Sept 2026
- paperGet To The Point: Summarization with Pointer-Generator Networks, arXiv · read 28 Sept 2026
- paperOn Faithfulness and Factuality in Abstractive Summarization, arXiv · read 28 Sept 2026
- paperROUGE: A Package for Automatic Evaluation of Summaries, Association for Computational Linguistics · read 28 Sept 2026